A text correction method, device, equipment and program product are described

CN122819230APending Publication Date: 2026-09-25SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610661400.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本申请提供一种描述文本纠错方法、装置、设备及程序产品,用于解决对多个VLM生成的描述文本进行融合或选择得到的最终的描述文本质量较差的问题

Benefits of technology

[0034]本申请在上述各方面提供的实现方式的基础上,还可以进行进一步组合以提供更多实现方式。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819230A_ABST
    Figure CN122819230A_ABST
Patent Text Reader

Abstract

A text correction method, device, equipment and program product are disclosed, and relate to the technical field of artificial intelligence generated content. After obtaining a plurality of description texts generated by a plurality of VLMs, the plurality of initial description texts are corrected based on the semantic differences between the plurality of initial description texts. Since the semantic differences between the plurality of initial description texts are considered, the plurality of target description texts after correction can avoid differences and description bias, improving the accuracy of the plurality of target description texts for video description. Based on the plurality of target description texts, text semantic fusion is performed to obtain a fused description text, improving the quality of the final generated fused description text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence generated content (AIGC) technology, and in particular to a method, apparatus, device and program product for correcting caption errors. Background Technology

[0002] In recent years, with the rapid development of artificial intelligence technology, the visual language model (VLM) has demonstrated powerful capabilities in the field of multimodal understanding. VLM can process image and video content and generate corresponding descriptive text.

[0003] In practical applications, a single VLM itself is biased when generating descriptive text for a video. Therefore, multiple VLMs are used to generate descriptive text for the same video. Then, the generated descriptive texts are merged or selected through simple rules (such as selecting the most frequently occurring description, selecting a description of appropriate length, or manually selecting the best description) to obtain the final descriptive text.

[0004] Because the description texts generated by multiple VLMs have different biases, the above-mentioned rule-based approach cannot correct this systematic bias, resulting in poor quality of the final generated description text. Summary of the Invention

[0005] This application provides a descriptive text correction method, apparatus, device, and program product to solve the problem of poor quality of the final descriptive text obtained by fusing or selecting descriptive texts generated by multiple VLMs.

[0006] Firstly, a method for correcting descriptive text is provided. This method includes: acquiring multiple initial descriptive texts generated by multiple Video Modules (VLMs) for a video; correcting the multiple initial descriptive texts based on semantic differences between them to obtain multiple corrected target descriptive texts; and performing semantic fusion on the multiple target descriptive texts to obtain a fused descriptive text.

[0007] Based on the aforementioned descriptive text correction method, after obtaining multiple descriptive texts generated by multiple VLMs, errors are corrected on these initial descriptive texts based on their semantic differences. By considering these semantic differences, the corrected target descriptive texts avoid discrepancies and descriptive biases, thus improving the accuracy of the video descriptions. Furthermore, textual semantic fusion is performed on the multiple target descriptive texts to obtain the fused descriptive text, further improving the quality of the final generated fused descriptive text.

[0008] One possible implementation involves performing at least one round of error correction iterations on multiple initial description texts based on semantic differences, resulting in multiple corrected target description texts. Specifically, in the first round of error correction iterations, the initial description texts are corrected based on semantic differences, resulting in multiple intermediate description texts for the first round of error correction iterations. In the m-th round of error correction iterations, the intermediate description texts obtained in the (m-1)-th round of error correction iterations are corrected based on semantic differences, resulting in multiple intermediate description texts for the m-th round of error correction iterations, where m is a positive integer greater than or equal to 2. The multiple target description texts are the multiple intermediate description texts obtained in the last round of error correction iterations when the error correction iteration stopping condition is met. Thus, the process of generating target description texts includes at least one round of error correction iterations, with each round of error correction involving intermediate description texts from the previous round, thereby progressively improving the quality of the intermediate description texts through at least one round of error correction iterations.

[0009] As one possible implementation, the stopping condition for error correction iteration is: the semantic difference rate of the multiple intermediate description texts obtained in the current round of error correction iteration is less than the semantic difference rate threshold; or the number of error correction iteration rounds reaches the round number threshold.

[0010] Thus, if the semantic difference rate of the multiple intermediate description texts obtained in the current round of error correction iteration is less than the semantic difference rate threshold, it can be ensured that the semantic differences of the multiple intermediate description texts obtained in the current round of error correction iteration are basically eliminated. The accuracy of the multiple intermediate description texts is improved and the descriptive bias is reduced, thereby improving the quality of the multiple intermediate description texts. By stopping the error correction iteration after the number of error correction iterations reaches the number of iterations threshold, the problem of being unable to correct the description text to meet the semantic difference rate requirement and getting stuck in an endless error correction iteration loop can be avoided, thereby improving the usability of the error correction iteration scheme.

[0011] As one possible implementation, in any round of error correction iteration, semantic difference analysis is performed on multiple descriptive texts to obtain the difference results corresponding to this round of error correction iteration. The difference results are used to indicate the semantic differences between every two descriptive texts. Error correction is performed based on the difference results corresponding to this round of error correction iteration and the multiple descriptive texts to obtain multiple intermediate descriptive texts for this round of error correction iteration. Specifically, in the first round of error correction iteration, the multiple descriptive texts are multiple initial descriptive texts, and in the m-th round of error correction iteration, the multiple descriptive texts are multiple intermediate descriptive texts obtained in the (m-1)th round of error correction iteration. Since the semantic differences between each descriptive text and other descriptive texts are determined in each round of error correction iteration, the multiple intermediate descriptive texts obtained in each round of error correction iteration can avoid differences and descriptive biases, thereby improving the accuracy of each intermediate descriptive text in describing the video.

[0012] One possible approach is to input multiple descriptive texts into a difference analysis model to obtain the difference results between every two descriptive texts in the current error correction iteration. In this way, difference analysis is performed uniformly on multiple descriptive texts through the difference analysis model in any error correction iteration, thereby improving the efficiency of difference analysis and avoiding resource waste.

[0013] As one possible implementation, multiple Virtual Machines (VLMs) include a first VLM and a second VLM. The first and second descriptive texts are input into the first VLM for semantic difference analysis, yielding the first difference result corresponding to the current error correction iteration. The first difference result indicates the semantic difference between the first and second descriptive texts, and includes the first difference result itself. In the first error correction iteration, the first and second descriptive texts are the initial descriptive texts generated for the video by the first and second VLMs, respectively. In the m-th error correction iteration, the first and second descriptive texts are the intermediate descriptive texts obtained from the (m-1)-th error correction iteration. Based on the first VLM and the first difference result, the first descriptive text is corrected to obtain the first intermediate descriptive text corresponding to the current error correction iteration.

[0014] Thus, by having any VLM independently perform difference analysis on its own generated description text compared to the description text generated by other VLMs, interference from other description texts in the difference analysis of the first and second description texts is avoided, thereby improving the accuracy of the generated first difference result. Furthermore, by correcting the description text generated by any VLM, the corrected intermediate description text output by the VLM corresponding to this round of correction iteration is obtained, thereby improving the accuracy of VLM error correction.

[0015] As one possible implementation, a first correction prompt is generated based on the first difference result. This first correction prompt guides the first Virtual Model (VLM) to correct the first descriptive text. The video, the first descriptive text, and the first correction prompt are input into the first VLM. The first VLM then corrects the first descriptive text, resulting in the first intermediate descriptive text output by the first VLM in this correction iteration. Thus, by generating the first correction prompt corresponding to the first difference result and inputting it into the first VLM, the model can be guided to accurately correct the first difference result, improving the efficiency and accuracy of model correction.

[0016] As one possible implementation, the first difference result is used to indicate different levels and types of semantic differences between the first descriptive text and the second descriptive text. The levels include at least one of the following: global, object, scene, target, or action in the video. The types include at least one of the following: replacement difference, omission difference, and addition difference. Replacement differences are used to indicate semantic differences that are different between the first description text and the second description text. Omission differences are used to indicate semantic differences that are not described in the first description text but are described in the second description text. Added differences are used to indicate semantic differences that are described in the first description text but not in the second description text.

[0017] Based on the above implementation method, by classifying the difference results, it is possible to more accurately identify each semantic difference and then perform accurate error correction, thereby improving the efficiency and accuracy of error correction.

[0018] As one possible implementation, the first difference result includes at least one difference item, each of which indicates a semantic difference content of a certain type at a certain level. A first error correction prompt word is generated for each difference item. Thus, a corresponding first error correction prompt word is generated for each difference item in the first difference result, enabling individual error correction for each difference item and improving the accuracy of error correction.

[0019] As one possible implementation, in the dialogue corresponding to each difference item in the first VLM, the input includes a video, a first descriptive text, each difference item, and a corresponding first correction prompt. Based on the video and the first correction prompt corresponding to each difference item, the first VLM corrects at least one difference item in the first descriptive text, obtaining the first intermediate descriptive text output by the first VLM in this round of correction iteration. In this way, the first VLM corrects each difference item separately in a single dialogue, avoiding interference from other differences items, thereby improving the accuracy of correction.

[0020] One possible implementation involves cropping video segments corresponding to the time ranges of target differences at the local level within the video. The local level refers to all levels except the global level. In the dialogue corresponding to the target difference in the first Video Library (VLM), the video segment, the first descriptive text, the target difference, and the corresponding first correction prompt are input. Thus, for target differences at the local level, since the target difference only involves a portion of the video content, only the video segment needs to be input for correction in the first VLM, rather than the entire video. This reduces the video memory resources consumed in correcting target-level differences and improves resource utilization.

[0021] As one possible implementation, the first VLM corrects the target differences in the first descriptive text based on the first correction prompt word corresponding to the target difference item and the video clip. Thus, for target differences at the local level, since the target difference item only involves a portion of the video content, the first VLM corrects based on the video clip. The video clip has fewer video frames than the complete video and is more targeted, thereby improving the efficiency and accuracy of error correction.

[0022] As one possible implementation, based on the target number of sampled frames per unit duration, frames are extracted from the video segment using a first Video Module (VLM) to obtain target video frames. The target number of sampled frames is greater than the number of sampled frames per unit duration extracted by the first VLM from the video. When extracting frames from a video segment, the target number of sampled frames per unit duration is larger than the number of sampled frames per unit duration extracted from the video, thus enabling the acquisition of more video frames from the video segment. Based on the first VLM, the target difference item in the first descriptive text is corrected according to the first error correction prompt word corresponding to the target difference item and the target video frame. Thus, compared to extracting frames using the number of sampled frames per unit duration, error correction is performed by extracting a larger number of video frames from the video segment based on the target number of sampled frames, thereby achieving refined error correction of the target difference item and improving the accuracy of error correction.

[0023] One possible implementation involves inputting the video, the first descriptive text, at least one difference item, the corresponding first error correction prompt, and the time range corresponding to the difference items at the local level in the video into a single round of dialogue in the first VLM. The local level refers to all levels other than the global level. In this way, all difference items in the first difference result between the first and second descriptive texts are simultaneously corrected in a single round of dialogue in the first VLM, thereby improving the efficiency of error correction.

[0024] One possible approach is to input multiple target descriptive texts into a large language model (LLM), perform semantic fusion based on the LLM, and obtain the fused descriptive text.

[0025] As one possible approach, multiple target description texts are semantically fused using a weighted fusion algorithm to obtain the fused description text.

[0026] Based on the two implementation methods mentioned above, the semantic integrity and quality of the final fused descriptive text are improved by semantically fusing multiple descriptive texts.

[0027] In a second aspect, a descriptive text correction apparatus is provided, the apparatus comprising modules for performing the descriptive text correction method in the first aspect or any possible implementation thereof.

[0028] The text correction device described in the second aspect can be a computing device, or a chip (system), network card or other component or assembly that can be set in a computing device, or a device that includes a computing device. This application does not limit it in this regard.

[0029] Furthermore, the technical effects of the text correction device described in the second aspect can be referred to the technical effects of the text correction method described in the first aspect, and will not be repeated here.

[0030] Thirdly, a computing device cluster is provided, the computing device cluster including at least one computing device, each computing device including a processor and a memory; the processor is coupled to the memory; the memory is used to store computer instructions, which are loaded and executed by the processor to enable the computing device cluster to perform the operational steps of the method described in any possible implementation of the first aspect above.

[0031] Fourthly, a computer-readable storage medium is provided, comprising: computer software instructions; when the computer software instructions are executed in a computer, causing the computer to perform operational steps of the method as described in any possible implementation of the first aspect.

[0032] Fifthly, embodiments of this application provide a chip system. The chip system includes a memory and at least one processor. The memory stores a set of computer instructions, which, when executed by the processor, perform the operational steps of the method described in any possible implementation of the first aspect.

[0033] Sixthly, a computer program product is provided that, when run on a computer, causes the computer to perform the operational steps of the method as described in any possible implementation of the first aspect.

[0034] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description

[0035] Figure 1 This is a schematic diagram illustrating a process for generating descriptive text for video using a single VLM. Figure 2 This is a schematic diagram illustrating the relationship between the number of model parameters and video memory usage in VLM. Figure 3 A schematic diagram illustrating the process of multi-VLM video description fusion scheme; Figure 4 This application provides a schematic diagram illustrating the architecture of a possible text correction system. Figure 5 A schematic diagram of the architecture of a cloud environment provided in an embodiment of this application; Figure 6 This application provides a schematic diagram of the system architecture for a text correction system 600 when used as a cloud service, as described in an embodiment of the present application. Figure 7 A flowchart illustrating a text correction method provided in an embodiment of this application; Figure 8 This is a schematic diagram illustrating the exchange of descriptive text among three VLMs, as provided in an embodiment of this application. Figure 9 A schematic diagram illustrating text semantic difference analysis for each VLM provided in the embodiments of this application; Figure 10 A schematic diagram of a generated error correction task group provided in an embodiment of this application; Figure 11 This is a schematic diagram of an execution error correction task group provided in an embodiment of this application; Figure 12 This is a schematic diagram illustrating a refined frame extraction of a video segment corresponding to a target difference item, provided as an embodiment of this application. Figure 13 A schematic diagram of video frames showing a robot arm operating a vegetable shelf, provided in an embodiment of this application; Figure 14 A schematic diagram illustrating a text correction device 1400 provided for an embodiment of this application; Figure 15 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Figure 16 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application; Figure 17 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application; Figure 18 This is a schematic diagram of a network connection structure between computing devices provided in an embodiment of this application. Detailed Implementation

[0036] This application's embodiments can be applied not only to existing video semantic understanding scenarios, and virtual environment understanding scenarios such as metaverse, virtual reality (VR), or augmented reality (AR), but also to future scenarios such as understanding two-dimensional or three-dimensional images. This application's embodiments do not impose any limitations on these applications. The fields in which this application's embodiments can be applied include autonomous driving, medical image diagnostic assistance, education, and industrial quality inspection. The solutions provided in this application's embodiments can be deployed as services on public, private, or hybrid cloud platforms, or they can run as algorithms on electronic devices or be integrated into device hardware. They can also be provided as independent software libraries or software toolkits. This application's embodiments do not impose any limitations on these applications.

[0037] To facilitate understanding, the relevant terms involved in the embodiments of this application will be introduced below.

[0038] (1) VLM VLM refers to a large-scale artificial intelligence model that processes both visual and linguistic information simultaneously.

[0039] (2) Description text A captipn is a sentence or paragraph generated by VLM that uses natural language to describe the content of an image.

[0040] (3) Frame skipping Frame extraction refers to the technique of extracting a portion of video frames from a video.

[0041] To better understand the issue of poor quality in the final description text obtained by fusing or selecting description texts generated by multiple VLMs in the background technology described above, the following section will combine... Figures 1 to 3 The paper describes a scheme for generating descriptive text using a single VLM and a scheme for fusing video descriptions using multiple VLMs.

[0042] In recent years, with the rapid development of artificial intelligence technology, Video Modeling (VLM) has demonstrated powerful capabilities in the field of multimodal understanding. VLM can process image and video content and generate corresponding descriptive text. It plays a crucial role in the data processing workflow of video semantic understanding, intelligent question answering, and artificial intelligence generated content (AIGC). Especially during the training of video generation models, high-quality descriptive text is key to improving model performance.

[0043] Figure 1 This is a schematic diagram illustrating a process for generating descriptive text for video using a single VLM.

[0044] like Figure 1As shown, firstly, the electronic device preprocesses the video, including keyframe extraction and image sampling, thereby transforming the continuous video stream into a series of discrete, model-processable video frames, which can be called a video frame sequence. Then, the electronic device inputs the video frame sequence into a pre-trained Visual Learning Model (VLM). Based on the VLM's vast internal parameters and learned visual-language association knowledge, it understands and analyzes the input video frame sequence, ultimately outputting a descriptive text describing the video content.

[0045] However, in practical applications, the scheme of generating high-quality descriptive text for videos using a single VLM faces significant technical problems, including limitations in computing resources and inherent biases in the model.

[0046] The limitation of computational resources refers to the constraint imposed by GPUs, video memory, and other computing resources. When generating descriptive text for video using a single Video Modeling Library (VLM), a balance must be struck between the size of the model's parameters and the number of video frames sampled / image precision. On one hand, while a VLM with a large number of parameters has stronger interpretation capabilities, its massive memory consumption necessitates reducing the number of frames sampled per second or the number of image sampling pixels. This can lead to missed keyframes or unclear video frames, affecting the accuracy of the descriptive text. On the other hand, while a VLM with fewer parameters can process more video frames or obtain clearer frames, its inherent limitations can result in recognition errors, missed details, or model illusions, leading to inaccurate descriptive text.

[0047] To alleviate the demand for video memory resources by VLM (Virtual Model), existing techniques employ quantization of VLM, which reduces the video memory usage during model runtime by decreasing the numerical precision of model parameters. For example, quantization is performed from 16-bit floating-point (FP16) / 16-bit brain floating-point (BFP16) format to 8-bit integer (INT8) or 4-bit integer (INT4) formats.

[0048] Figure 2 This diagram illustrates the relationship between the number of model parameters and video memory usage in a VLM (Virtual Machine Model).

[0049] like Figure 2As shown, with model parameter precisions of FP16 / BFP16, INT8, and INT4, the memory usage increases with the number of model parameters. Furthermore, for the same number of model parameters, the memory usage is highest when the model parameters are in FP16 / BFP16 floating-point format and lowest when they are in INT4 integer format; that is, the memory usage decreases as the model parameter precision decreases. Therefore, under the same hardware resource conditions, reducing the model parameter precision can support the deployment or operation of VLMs with a larger number of parameters.

[0050] While the aforementioned quantization techniques can alleviate the demand for video memory by VLM to some extent, they come at the cost of model accuracy and performance. The quantization process introduces errors, which may lead to a decrease in the quality of the descriptive text generated by VLM.

[0051] The inherent bias of a model means that even VLMs with a large number of parameters produce results that reflect the model's own biases. For example, some VLMs may be more inclined to describe static scene backgrounds in detail while ignoring dynamic behaviors, while others may be more inclined to describe dynamic behaviors in detail while ignoring static scene backgrounds. This inherent bias of the model can lead to a situation where the descriptive text generated by a single VLM cannot comprehensively and evenly reflect the video content.

[0052] To address the issue of model capability bias, the traditional approach is to employ a multi-VLM video description fusion scheme.

[0053] Figure 3 This is a schematic diagram illustrating the process of a multi-VLM video description fusion scheme.

[0054] like Figure 3 As shown, after video frame extraction and sampling (i.e., keyframe extraction and image sampling) to obtain a video frame sequence, these sequences are input into VLMA, VLMB, and VLMC, respectively. VLMA generates video description text A, VLMB generates video description text B, and VLMC generates video description text C. Description texts A, B, and C are then simply fused or selected based on a voting mechanism or rule-based selection to obtain the final description text output.

[0055] However, the above approach still has significant drawbacks. First, it lacks a systematic analysis of the deep semantic differences between multiple descriptive texts. If multiple VLMs contain errors or biases on the same detail, a simple voting mechanism cannot detect and correct such systematic biases.

[0056] Secondly, it cannot handle complex semantic conflicts. For example, for the same video, the descriptive text output by one model may describe action A in the video, while the descriptive text output by another model may describe action B in the video. Simply choosing one may ignore the other action that actually occurred.

[0057] Most importantly, the multi-VLM video description fusion scheme is a passive selection rather than an active error correction. When the description texts generated by all VLMs are unsatisfactory, the system cannot generate new description texts of higher quality. Therefore, the description texts generated by the above method are of poor quality.

[0058] To address the technical problem of poor quality of the final description text generated by the aforementioned multi-VLM video description fusion scheme, this application provides a description text correction method, apparatus, device, and program product. In particular, it provides a method that corrects errors in multiple initial description texts based on their semantic differences, and then performs text semantic fusion to obtain the final fused description text. After obtaining multiple description texts generated by multiple VLMs, errors are corrected based on the semantic differences between the initial description texts. By considering the semantic differences between the initial description texts, discrepancies and descriptive biases are avoided in the corrected target description texts, improving the accuracy of the video descriptions. Then, text semantic fusion is performed based on the multiple target description texts to obtain the fused description text, further improving the quality of the final generated fused description text.

[0059] Next, combine Figures 4 to 13 The implementation of the embodiments of this application will be described.

[0060] Figure 4 This is a schematic diagram illustrating the architecture of a possible text correction system provided in an embodiment of this application.

[0061] like Figure 4 As shown, the description text error correction system 400 may include multiple VLMs, a VLM interaction module 410, a description text exchange and analysis module 420, an error correction task generation module 430, an error correction task execution module 440, a description text fusion module 450, and a video capture module 460.

[0062] In the descriptive text correction system 400, during the descriptive text correction process, the VLM interaction module 410 and the descriptive text exchange and analysis module 420 are communicatively connected. The descriptive text exchange and analysis module 420 is communicatively connected to the error correction task generation module 430 and the video capture module 460. The error correction task generation module 430 is communicatively connected to the error correction task execution module 440. The error correction task execution module 440 is communicatively connected to the descriptive text fusion module 450 and the video capture module 460.

[0063] The VLM interaction module 410 is used to interact with multiple VLMs, obtain multiple initial description texts generated by multiple VLM models for the video, and send them to the description text exchange and analysis module 420.

[0064] The video can be from a real-world scene, such as video captured by a camera or a live video stream. It can also be from a virtual scene, such as video captured in virtual environments like the metaverse, VR, or AR.

[0065] Videos can be from different fields. For example, videos could be in-vehicle videos in the field of autonomous driving, medical videos in the field of medical diagnosis, course videos or experiment videos in the field of education, or production line inspection videos in the field of industrial quality inspection, etc.

[0066] The descriptive text exchange and analysis module 420, after acquiring multiple initial descriptive texts, performs semantic difference analysis on these texts to obtain the semantic differences between them, and sends this result to the error correction task generation module 430. The descriptive text exchange and analysis module 420 also determines local difference time periods and sends these local difference time periods to the video capture module 460. The local difference time period refers to the time range corresponding to the semantic differences at a local level in the difference results.

[0067] The video extraction module 460 extracts video segments based on local difference time periods and sends these segments to the error correction task module 440. The extraction process maintains the original video format and bitrate, manipulating only the video length.

[0068] The error correction task generation module 430 is used to generate error correction tasks based on the semantic differences between multiple initial description texts, videos, and video clips.

[0069] The error correction task module 440 performs error correction tasks to correct multiple initial descriptive texts, resulting in multiple corrected target descriptive texts. The error correction task module 440 then sends these multiple target descriptive texts to the descriptive text fusion module 450.

[0070] The description text fusion module 450 performs text semantic fusion on multiple target description texts to obtain the fused description text.

[0071] The text correction system 400 described above can also be hardware such as electronic devices, computing device clusters, data centers, or edge devices, or it can be a functional module deployed in hardware such as electronic devices, computing device clusters, data centers, or edge devices. This application embodiment does not limit this.

[0072] The descriptive text correction system 400 can also be a software system. Specifically, a software system can be a software service, algorithm, application development tool, toolchain, etc. For example, when the software system is a software service, the descriptive text correction system can be a cloud service software product with a web interface, a software product integrated into the functionality of a desktop software entity, a data processing workflow software, or a software product in the form of an API interface that only exposes external interaction, etc. This application embodiment does not impose any limitations on this. Wherein, when the descriptive text correction system 400 is a software service, it can be sold in the form of a subscription for model usage rights, or by pricing the number of tokens consumed by model inference.

[0073] For example, the text correction system 400 is deployed in Figure 5 When a cloud service is provided on a cloud platform in the cloud environment shown, the cloud environment can be a public cloud, a private cloud, or a hybrid cloud, etc., and this application embodiment does not limit this.

[0074] For example, in scenarios where software as a service (SaaS) is provided to small and medium-sized enterprises or individuals, the descriptive text correction system 400, deployed in a public cloud, can adopt a multi-tenant architecture, meaning that each tenant's descriptive text correction system 400 is isolated from each other. The descriptive text correction system 400 can dynamically allocate computing power and encrypt data to ensure data security.

[0075] In scenarios where services are provided to large, classified organizations, to ensure data privacy, the text correction system 400 is deployed in a private cloud. That is, the text correction system is deployed in a localized data center and interfaces with the organization's internal systems to ensure that data is not leaked externally.

[0076] In scenarios where services are provided to medium-sized enterprises, to balance cost and data security, the Text Correction System 400 is deployed in a hybrid cloud. Specifically, the Text Correction System 400 is deployed across multiple Virtual Machines (VLMs): VLMs containing core data are deployed in a private cloud, while VLMs containing non-core data are deployed in a public cloud, with an added inter-cloud synchronization module.

[0077] Figure 5 This is a schematic diagram of the architecture of a cloud environment provided in an embodiment of this application.

[0078] like Figure 5 As shown, cloud platform 510 is located in cloud data center 520, and cloud platform 510 is connected to client 530 via the Internet. Cloud platform 510 is used to manage at least one server in cloud data center 520.

[0079] like Figure 5As shown, cloud platform 510 is connected to server 1 and server 2 via the internal network of the data center. Cloud platform 510 is used to manage server 1 and server 2 of cloud data center 520. Figure 5 Only two servers are shown, but it is not limited to two servers. Figure 5 Only one cloud data center is shown, but it is not limited to one cloud data center.

[0080] The functions provided by Cloud Platform 510 include: providing access interfaces (such as user interfaces or application programming interfaces (APIs)). Tenants can operate clients to remotely access the access interface to register a cloud account and password on the cloud platform and log in to the cloud platform. After the cloud platform successfully authenticates the cloud account and password, the tenant can further pay to select and purchase virtual machines of specific specifications (processor, memory, disk) on the cloud platform. After successful payment and purchase, the cloud platform provides the remote login account and password for the purchased virtual machine. The client can remotely log in to the virtual machine and install and run the tenant's application in the virtual machine.

[0081] The cloud platform's logical functional divisions are as follows: User Console, Compute Management Service, Network Management Service, Storage Management Service, Authentication Service, and Image Management Service. The User Console provides an interface or API for interaction with tenants. The Compute Management Service manages servers running virtual machines and containers, as well as bare metal servers. The Network Management Service manages network services (such as gateways and firewalls). The Storage Management Service manages storage services (such as data bucket services). The Authentication Service manages tenant account passwords. The Image Management Service manages virtual machine images.

[0082] The hardware layer of each server in Server 1 and Server 2 includes memory, network interface card, disk, bus, and processor.

[0083] The memory may include volatile memory, such as random access memory (RAM), or non-volatile memory, such as read-only memory (ROM) or flash memory. The disk may include hard disk drives (HDDs) or solid-state drives (SSDs).

[0084] Network cards can include Ethernet cards, wireless cards, modem cards, or fiber optic cards, etc.

[0085] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 5 The bus is represented by a single line, but this does not mean that there is only one bus or one type of bus. A bus can include a path for transferring information between various components of the server (e.g., memory, network card, disk, processor).

[0086] The processor can be a neural processing unit (NPU), a tensor processing unit (TPU), an intelligence processing unit (IPU), a data processing unit (DPU), a field programmable gate array (FPGA), or an application-specific integrated circuit (ASIC).

[0087] Servers 1 and 2 are used to generate and allocate computing resources according to user needs based on virtualization technology. In this embodiment, the computing resources run some or all of the hardware resources of virtual devices, which are used by allocating the computing resources to virtual devices running on the servers. The virtual devices include virtual machines or containers.

[0088] like Figure 5 As shown, the software layer of each server in Server 1 and Server 2 includes a host operating system and virtual machines. The host operating system includes a virtual machine manager, which includes a cloud platform client. The software layer of Server 1 includes Virtual Machine 1 and Virtual Machine 2, and the software layer of Server 2 includes Virtual Machine 3 and Virtual Machine 4.

[0089] The above text Figure 5 The diagram shows the overall architecture of the cloud environment. Next, we will combine... Figure 6 The following describes a specific system architecture example of the text correction system 600 provided in the embodiments of this application when it is used as a cloud service on a cloud platform 510.

[0090] Figure 6This is a schematic diagram of the system architecture of a text correction system 600 as a cloud service, provided as an embodiment of this application.

[0091] like Figure 6 As shown, end user 610 accesses cloud service 620 provided by cloud service provider via the Internet. Cloud service 620 includes user operation and interface service 621, process control and task processing module 622, video basic processing unit 623, video description inference processing unit 624, and video description fusion processing unit 625.

[0092] The descriptive text correction system 600 can be composed of a process control and task processing module 622, a video basic processing unit 623, a video description inference processing unit 624, and a video description fusion processing unit 625. The process control and task processing module 622 is equivalent to... Figure 4 The VLM interaction module 410, description text exchange and analysis module 420, error correction task generation module 430, and error correction task execution module 440 are included. The video basic processing unit 623 is equivalent to... Figure 4 The video capture module 460 in the middle. The video description inference processing unit 624 is equivalent to Figure 4 The video description fusion processing unit 625 is a unit composed of multiple VLMs. Figure 4 The descriptive text fusion module 450 in the middle.

[0093] In the text correction system 600, during the text correction process, the user operation and interface service 621 is communicatively connected to the flow control and task processing module 622. The flow control and task processing module 622 is communicatively connected to the video basic processing unit 623, the video description inference processing unit 624, and the video description fusion processing unit 625. The video basic processing unit 623 includes a video format processor and a video cropping processor. The video description inference processing unit 624 includes multiple VLMs. Figure 6 The diagram shows n models: VLM 1, VLM 2, ..., VLM n. The video description fusion processing unit 624 includes a large language model.

[0094] End user 610 invokes user operation and interface service 621 to send video and video description requirements to cloud service 620 via the Internet. The video description requirements include specifications for the scope, text format, and level of detail of the video description.

[0095] The cloud service 620 calls the video basic processing unit 623 through the process control and task processing module 622 to convert the video content into a video format that the VLM model can process.

[0096] The process control and task processing module 622 calls the video description inference processing unit 624, where multiple VLMs (Video Description Models) generate multiple initial description texts for the format-converted video or multiple intermediate description texts after error correction. These multiple VLMs can be deployed across multiple cloud service instances, each with its own computing resources and inference framework. Since the inference efficiency and response speed of each VLM differ, the process control and task processing module 622 schedules the requests to be appropriately allocated to each cloud service instance.

[0097] The process control and task processing module 622 corrects the multiple descriptive texts in each round of error correction iteration, resulting in multiple corrected target descriptive texts. These multiple target descriptive texts can also be output to the end user 610, who can then decide whether to perform fusion processing.

[0098] The process control and task processing module 622 calls the video description fusion processing unit 625, where a large language model deployed in the video description fusion processing unit 625 performs text semantic fusion on multiple target description texts to obtain the fused description text. The cloud service 620 outputs the target description text to the end user 610.

[0099] The following is combined Figures 4-6 The content shown provides a detailed description of the text correction method provided in the embodiments of this application.

[0100] Figure 7 This is a flowchart illustrating a text correction method provided in an embodiment of this application. The following describes the process as follows: Figure 4 The description shown is provided by the text correction system 400. (About...) Figure 7 The description of the text correction system 400 can be found in the above text. Figure 4 The relevant descriptions in the original text will not be repeated in the embodiments of this application.

[0101] like Figure 7 As shown, the process may include steps 701 to 703.

[0102] Step 701: The description text correction system 400 obtains multiple initial description texts generated by multiple VLMs for the video.

[0103] The VLM can be a general VLM or a special VLM for different fields. This application does not limit this.

[0104] When the video is in-vehicle video, the Visual Management System (VLM) can be a VLM specifically designed for autonomous driving, with descriptive text describing the scene or event, thus enabling its use in accident backtracking and autonomous driving algorithm optimization. When the video is medical video, the VLM can be a medical-specific VLM, with descriptive text describing the location of lesions, thus assisting doctors in lesion localization and reducing missed or misdiagnosed cases. When the video is course video, the VLM can be a general-purpose VLM, with descriptive text describing the course content, thus enabling its use in knowledge point extraction and retrieval. When the video is production line inspection video, the VLM can be an industrial-specific VLM, with descriptive text describing product defects and quality inspection judgments, thus assisting in the identification of product defects.

[0105] As one possible implementation, the descriptive text correction system 400 can integrate multiple VLMs. After acquiring the video, the descriptive text correction system 400 inputs it into multiple VLMs, performs video description inference based on the multiple VLMs, and obtains the corresponding descriptive text generated by each VLM.

[0106] Optionally, the description text correction system 400 can provide selection options for VLMs through a visual interface, and determine multiple VLMs for video description inference based on the user's selection instructions.

[0107] As one possible implementation, the descriptive text correction system 400 can receive multiple initial descriptive texts from the inference device. The inference device integrates multiple Video Description Models (VLMs), and uses these VLMs to perform video description inference on the same video, obtaining multiple initial descriptive texts. The inference device then sends these multiple initial descriptive texts to the descriptive text correction system 400.

[0108] The descriptive text correction system 400 may integrate an error correction model or integrate multiple VLMs identical to those in the inference device. The error correction model or multiple VLMs in the descriptive text correction system 400 are used to correct multiple initial descriptive texts.

[0109] Step 702: The description text correction system 400 corrects the multiple initial description texts based on the semantic differences between them, and obtains multiple target description texts after correction.

[0110] As one possible implementation, the descriptive text correction system 400 can perform at least one round of error correction iteration on multiple initial descriptive texts based on the semantic differences between them, to obtain multiple target descriptive texts after error correction.

[0111] Specifically, in the first round of error correction iteration, the description text correction system 400 corrects multiple initial description texts based on the semantic differences between them, thus obtaining multiple intermediate description texts for the first round of error correction iteration.

[0112] In the m-th round of error correction iteration, the description text correction system 400 corrects the multiple intermediate description texts obtained in the (m-1)-th round of error correction iteration based on the semantic differences between the multiple intermediate description texts obtained in the (m-1)-th round of error correction iteration, thus obtaining multiple intermediate description texts in the m-th round of error correction iteration, where m is a positive integer greater than or equal to 2 and m is less than or equal to the total number of rounds of error correction iteration.

[0113] Here, multiple target description texts are multiple intermediate description texts obtained in the last round of error correction iteration when the error correction iteration stopping condition is met. The error correction iteration stopping condition is: the semantic difference rate corresponding to the multiple intermediate description texts obtained in the current round of error correction iteration is less than the semantic difference rate threshold; or the number of error correction iteration rounds reaches the round number threshold.

[0114] Specifically, the descriptive text correction system 400 corrects multiple descriptive texts corresponding to each correction iteration based on the difference results generated in each correction iteration, resulting in multiple intermediate descriptive texts after correction. The descriptive text correction system 400 calculates the semantic difference rate between every two descriptive texts in the multiple intermediate descriptive texts after correction.

[0115] The descriptive text correction system 400 can be used when the semantic difference rate among multiple descriptive texts in the current round of correction iteration is less than the semantic difference rate threshold, or the average semantic difference rate is less than the semantic difference rate threshold, or there is a set number of semantic difference rates less than the semantic difference rate threshold. If the correction iteration has not reached the round number threshold, the next round of correction iteration will not be carried out, and the corrected intermediate descriptive texts obtained in the current round of correction iteration will be used as multiple target descriptive texts.

[0116] The description text correction system 400 can also continue to the next round of error correction iteration if the semantic difference rate corresponding to the current round of error correction iteration is greater than or equal to the semantic difference rate threshold, until the number of error correction iterations reaches the round number threshold and the semantic difference rate corresponding to the last round of error correction iteration is still greater than or equal to the semantic difference rate threshold, and obtain multiple intermediate description texts after error correction in the last round of error correction iteration as multiple target description texts.

[0117] The text correction system 400 does not restrict the error correction method in at least one round of error correction iteration; the error correction method in different rounds of error correction iteration can be the same or different. The specific implementation process of the error correction method in any round of error correction iteration can be found in the four optional implementation methods described below.

[0118] As a first optional implementation, the descriptive text correction system 400 can input multiple descriptive texts corresponding to the current error correction iteration into the error correction model in any error correction iteration, determine the semantic differences between the multiple descriptive texts based on the error correction model, and perform an error correction iteration based on the semantic differences to obtain multiple intermediate descriptive texts for the current error correction iteration.

[0119] The error correction model can be a single large language model, a deep learning model, or a combination of algorithms and models. This application does not limit this.

[0120] Specifically, in the first round of error correction iteration, multiple initial descriptive texts are input into the error correction model. The descriptive text error correction system 400 can determine the semantic differences between the multiple initial descriptive texts based on the error correction model, and perform one round of error correction iteration based on the semantic differences to obtain multiple intermediate descriptive texts for the first round of error correction iteration. The descriptive text error correction system 400 can use the multiple intermediate descriptive texts from the first round of error correction iteration as the multiple descriptive texts corresponding to the second round of error correction iteration.

[0121] The descriptive text correction system 400 can, in the m-th correction iteration, input multiple descriptive texts corresponding to the m-th correction iteration into the correction model to determine the semantic differences between the multiple descriptive texts, and perform one round of correction iteration based on the semantic differences to obtain multiple intermediate descriptive texts for the m-th correction iteration. The descriptive text correction system 400 can use the multiple intermediate descriptive texts for the m-th correction iteration as the multiple descriptive texts corresponding to the (m+1)-th correction iteration.

[0122] As a second optional implementation, the descriptive text correction system 400 performs semantic difference analysis on multiple descriptive texts in any round of correction iteration to obtain the difference result corresponding to this round of correction iteration. The descriptive text correction system 400 performs error correction based on the difference result corresponding to this round of correction iteration and multiple descriptive texts to obtain multiple intermediate descriptive texts after correction for this round of correction iteration.

[0123] In the first round of error correction iteration, the multiple description texts are multiple initial description texts. In the m-th round of error correction iteration, the multiple description texts are multiple intermediate description texts obtained in the (m-1)-th round of error correction iteration.

[0124] The difference results are used to indicate the semantic differences between any two descriptive texts. The difference results include the differences between any two descriptive texts across multiple descriptive texts. The difference results between any two descriptive texts refer to the semantic differences between one descriptive text and another.

[0125] For example, if multiple descriptive texts include descriptive text 1, descriptive text 2, descriptive text 3, and descriptive text 4, the difference results include: difference result 1 between descriptive text 1 and descriptive text 2, difference result 2 between descriptive text 1 and descriptive text 3, difference result 3 between descriptive text 1 and descriptive text 4, difference result 4 between descriptive text 2 and descriptive text 3, and difference result 5 between descriptive text 3 and descriptive text 4.

[0126] For example, the difference result 1 between description text 1 and description text 2 can include: the difference result between description text 2 and description text 1, with description text 2 as the subject, and the difference result between description text 1 and description text 2, with description text 1 as the subject.

[0127] In cases where multiple descriptive texts include a first descriptive text and a second descriptive text, the difference result between the first descriptive text and the second descriptive text includes a first difference result between the first descriptive text and the second descriptive text.

[0128] The first difference result specifically indicates different levels and types of semantic differences between the first descriptive text and the second descriptive text. Levels include at least one of the following: global, object, scene, target, or action in the video. Types include at least one of the following: replacement difference, omission difference, or addition difference.

[0129] Global refers to the environment, scene, and background in the video. Object refers to items, people, etc., in the video. Scene refers to location, lighting, weather, and time period in the video. Goal refers to the person or item that the video focuses on. Action refers to the activities of the people in the video. Replacement difference is used to indicate semantic differences between the first and second descriptive texts. Omission difference is used to indicate semantic differences that are not described in the first descriptive text but are described in the second descriptive text. Added difference is used to indicate semantic differences that are described in the first descriptive text but not in the second descriptive text.

[0130] For example, difference results can be divided into two levels and three types of differences. The two levels include the global level and the object level, and the three types include replacement differences, omission differences, and addition differences.

[0131] Table 1 is a difference table provided in the embodiments of this application.

[0132] Table 1

[0133] As shown in Table 1, the substitution difference at the global level refers to the difference between the description of the environment, scene, and background in the self-description text and other description texts. The omission difference at the global level refers to the absence of descriptions of the environment, scene, and background in the self-description text, while other description texts describe the missing content. The addition difference at the global level refers to the addition of content to the description of the environment, scene, and background in the self-description text compared to the descriptions in other description texts.

[0134] Object-level substitution differences refer to discrepancies between the self-description text and other descriptive texts regarding the name, appearance, and behavior of a specific item within a given timeframe. Object-level omission differences refer to a situation where the self-description text lacks descriptions of the name, appearance, and behavior of a specific item within a given timeframe, but these missing details are included in other descriptive texts. Object-level addition differences refer to situations where the self-description text includes additional descriptions of the name, appearance, and behavior of a specific item within a given timeframe compared to other descriptive texts.

[0135] In this application embodiment, there are no restrictions on the implementation method of performing text semantic difference analysis on multiple descriptive texts. The descriptive text error correction system 400 can perform text semantic difference analysis through a difference analysis model or VLM. For specific implementation process, please refer to the two possible implementation methods below.

[0136] As one possible implementation, the descriptive text correction system 400 can input multiple descriptive texts into a difference analysis model to obtain the difference results between every two descriptive texts in the current round of correction iterations.

[0137] The difference analysis model can be a large language model, a neural network model, etc., and this application does not limit it.

[0138] Specifically, after obtaining multiple descriptive texts corresponding to the current round of error correction iteration, the descriptive text correction system 400 can input the multiple descriptive texts into the difference analysis model. Based on the difference analysis model, the multiple descriptive texts are converted into vectors, and the difference results between each pair of descriptive texts are determined by calculating the similarity between the vectors.

[0139] For example, the descriptive text correction system 400 can add an intermediate coordination module between the VLM interaction module 410 and multiple VLMs. The intermediate coordination module deploys a large language model. The intermediate coordination module obtains multiple descriptive texts output by multiple VLMs and inputs them into the large language model. Based on the large language model, the intermediate coordination module converts every two descriptive texts into word vectors and determines the difference between each pair of descriptive texts by calculating the cosine distance between the corresponding word vectors of each pair of descriptive texts.

[0140] Optionally, the description text correction system 400 may, in the first round of error correction iteration, input multiple initial description texts generated by multiple VLMs for the video into a difference analysis model to obtain the difference results of every two initial description texts in the multiple initial description texts in the first round of error correction iteration.

[0141] The text correction system 400 can also input multiple intermediate descriptive texts obtained in the (m-1)th round of correction iterations into a difference analysis model to obtain the difference results in the m-th round of correction iterations. These difference results include the difference results between every two intermediate descriptive texts among the multiple intermediate descriptive texts obtained in the (m-1)th round of correction iterations.

[0142] The process of performing difference analysis based on the difference analysis model in each round of error correction iteration can be referred to the relevant description in the possible implementation methods above, and will not be repeated here.

[0143] As one possible implementation, multiple VLMs include a first VLM and a second VLM. The descriptive text correction system 400 inputs the first descriptive text and the second descriptive text into the first VLM for text semantic difference analysis, obtaining a first difference result output by the first VLM. The first difference result indicates the semantic difference between the first descriptive text and the second descriptive text, and the difference result includes the first difference result.

[0144] In the first round of error correction iteration, the first description text is the initial description text generated by the first VLM for the video, and the second description text is the initial description text generated by the second VLM for the video. In the m-th round of error correction iteration, the first and second description texts are the intermediate description texts obtained from the error correction in the (m-1)-th round of error correction iteration.

[0145] That is, in the m-th error correction iteration, the first description text and the second description text can be the intermediate description text obtained by the error correction model in the (m-1)-th error correction iteration. The first description text can be the intermediate description text obtained by the first VLM in the (m-1)-th error correction iteration, and the second description text can be the intermediate description text obtained by the second VLM in the (m-1)-th error correction iteration.

[0146] In the case of text semantic difference analysis using the VLM model, specifically, the descriptive text correction system 400 obtains the second descriptive text generated by the second VLM for the video and sends it to the first VLM. The descriptive text correction system 400 performs text semantic difference analysis based on the first VLM using its own generated first and second descriptive texts. This involves converting the first and second descriptive texts into vectors, calculating the similarity between the vectors, and outputting the first difference result based on the semantic difference between the first and second descriptive texts, determined by the first VLM.

[0147] For example, multiple VLMs include VLM1, VLM2, and VLM3. The descriptive text correction system 400 can also acquire descriptive text 1 generated by VLM1 for the video, descriptive text 2 generated by VLM2 for the video, and descriptive text 3 generated by VLM3 for the video. The descriptive text correction system 400 can also exchange descriptive texts among the three VLMs, ensuring that each VLM acquires the descriptive text generated by the other VLMs. Each VLM in the descriptive text correction system 400 performs text semantic difference analysis based on its own descriptive text and each descriptive text generated by each of the other VLMs, obtaining the difference results output by each VLM. The descriptive text correction system 400 can fill the difference results into a difference table as shown in Table 1, obtaining a structured difference result table.

[0148] Figure 8 This is a schematic diagram illustrating the exchange of descriptive text between three VLMs, as provided in an embodiment of this application.

[0149] like Figure 8 As shown, after VLM1 performs inference on the video, the result of VLM1 is descriptive text 1. After VLM2 performs inference on the video, the result of VLM2 is descriptive text 2. After VLM3 performs inference on the video, the result of VLM3 is descriptive text 3.

[0150] VLM1 obtains the results generated by inference from VLM2 and VLM3. VLM2 obtains the results generated by inference from VLM1 and VLM3. VLM3 obtains the results generated by inference from VLM1 and VLM2.

[0151] Figure 9 This is a schematic diagram illustrating text semantic difference analysis for each VLM provided in the embodiments of this application.

[0152] like Figure 9 As shown, the description text correction system 400, for VLM1, first inputs description text 1 and description text 2, analyzes the semantic differences between description text 1 and description text 2 to obtain the difference results; then inputs description text 1 and description text 3, analyzes the semantic differences between description text 1 and description text 3 to obtain the difference results.

[0153] The description text correction system 400, for VLM2, first inputs description text 2 and description text 1, analyzes the semantic differences between description text 2 and description text 1 to obtain the difference results; then inputs description text 2 and description text 3, analyzes the semantic differences between description text 2 and description text 3 to obtain the difference results.

[0154] The description text correction system 400, for VLM3, first inputs description text 3 and description text 1, analyzes the semantic differences between description text 3 and description text 1 to obtain the difference results; then inputs description text 3 and description text 2, analyzes the semantic differences between description text 3 and description text 2 to obtain the difference results.

[0155] Optionally, the descriptive text correction system 400 may, in the first round of error correction iteration, obtain the second initial descriptive text generated by the second VLM for the video and send it to the first VLM. The descriptive text correction system 400 performs semantic analysis based on the first initial descriptive text and the second initial descriptive text generated by the first VLM itself, to obtain the first difference result output by the first VLM in the first round of error correction iteration.

[0156] The descriptive text correction system 400 can also, in the m-th correction iteration, obtain the second intermediate descriptive text corrected by the second VLM in the (m-1)-th correction iteration and send it to the first VLM. Based on the first VLM performing semantic analysis on the first intermediate descriptive text corrected by the first VLM in the (m-1)-th correction iteration and the second intermediate descriptive text, the descriptive text correction system 400 obtains the first difference result output by the first VLM in the m-th correction iteration.

[0157] As one possible implementation, the text correction system 400 can also output the difference results generated in each round of correction iteration, allowing users to modify, supplement, and manually confirm the difference results to obtain updated difference results, thereby improving the accuracy of the difference results.

[0158] In this application embodiment, there is no limitation on the implementation method of correcting multiple description texts. The description text correction system 400 can correct multiple description texts through a correction model or multiple VLMs. For specific implementation process, please refer to the two possible implementation methods below.

[0159] As one possible implementation, the descriptive text correction system 400 can correct the first descriptive text based on the first difference result output by the difference analysis model or the first VLM based on the correction model, and obtain the first intermediate descriptive text corresponding to the current correction iteration.

[0160] As one possible implementation, the descriptive text correction system 400 can correct the first descriptive text based on the first VLM according to the difference analysis model or the first difference result output by the first VLM, to obtain the first intermediate descriptive text corresponding to the current correction iteration.

[0161] Optionally, the description text correction system 400 can input the first difference result, the first description text, and the video into the first VLM, and correct the first description text based on the first VLM to obtain the first intermediate description text corresponding to the current correction iteration.

[0162] The first intermediate description text can refer to the description text after the first VLM corrects the semantic differences in the first description text, or it can refer to the description text after the first VLM re-describes the video; the re-described description text achieves the correction of the semantic differences in the first description text.

[0163] Optionally, the descriptive text correction system 400 can generate a first correction prompt word based on the first difference result. The first correction prompt word is used to guide the first VLM to correct the first descriptive text. The descriptive text correction system 400 can input the video, the first descriptive text, and the first correction prompt word into the first VLM to correct the first descriptive text, and obtain the first intermediate descriptive text after correction in this round of correction iteration output by the first VLM.

[0164] Specifically, the descriptive text correction system 400 can generate a first correction prompt word for the first difference result based on the type of semantic difference content contained in the first difference result and different types of prompt word templates. Alternatively, the descriptive text correction system 400 can input the first difference result into a prompt word generation model, which then generates the first correction prompt word corresponding to the first difference result. The prompt word generation model can be a large language model, a deep learning model, etc., and this embodiment does not impose any limitations on it.

[0165] Specifically, when the first difference result includes semantic difference content of a replaced difference type, the first error correction prompt can be used to prompt the first VLM to verify the accuracy of the semantic difference content. When the first difference result includes semantic difference content of a newly added difference type, the first error correction prompt can be used by the first VLM to verify the accuracy of the semantic difference content; specifically, for semantic difference content described in the first description text but not in the second description text, it determines whether it exists in the video. When the first difference result includes semantic difference content of an omitted difference type, the first error correction prompt can be used to prompt the first VLM to verify the accuracy of the semantic difference content; specifically, for semantic difference content not described in the first description text but described in the second description text, it determines whether it exists in the video but is not described in the first description text.

[0166] For example, when the video is a medical video, the descriptive text correction system 400 can input the first difference result into a prompt word generation model, which then combines medical standards and standardized medical terminology to generate a first correction prompt word corresponding to the first difference result. When the video is an instructional video, the descriptive text correction system 400 can generate a first correction prompt word corresponding to the first difference result based on the difference type of the semantic difference content contained in the first difference result, as well as prompt word templates for different difference types and the teaching syllabus.

[0167] The process by which the descriptive text correction system 400 corrects errors in the first descriptive text based on the first VLM can be referred to below. Figures 10-12 The relevant descriptions are not repeated in the embodiments of this application.

[0168] Step 703: The description text correction system 400 performs text semantic fusion on multiple target description texts to obtain the fused description text.

[0169] As one possible implementation, the descriptive text correction system 400 can input multiple target descriptive texts into a large language model, perform semantic fusion based on the large language model, and obtain the fused descriptive text.

[0170] Specifically, the descriptive text correction system 400 can input multiple target descriptive texts into a large language model, and fuse the semantic content of the multiple target descriptive texts based on the large language model, ensuring that the semantics of the fused descriptive text includes the semantics of each descriptive text. The descriptive text correction system 400 can also input multiple target descriptive texts into the large language model, and based on the large language model, exclude descriptive texts whose semantic similarity to other descriptive texts is below a threshold, and select target descriptive texts whose semantic similarity to each other is greater than a threshold, and fuse them into a single target descriptive text.

[0171] As one possible implementation, the descriptive text correction system 400 can also perform semantic fusion of multiple target descriptive texts using a weighted fusion algorithm to obtain the fused descriptive text.

[0172] Weighted fusion algorithms are a common type of model fusion algorithm. They improve overall inference performance by weighting and averaging the outputs of different models. In this embodiment, semantic fusion using a weighted fusion algorithm involves weighting and averaging multiple target description texts to obtain the weighted average description text.

[0173] Specifically, the description text correction system 400 can use the weights corresponding to multiple VLMs as the weights of multiple target description texts, and perform weighted fusion of multiple target description texts according to the weights to obtain a final target description text.

[0174] The description text correction system 400 can also use semantic similarity to calculate the semantic similarity between each target description text and other target description texts in multiple target description texts, assign different weights to each target description text according to the semantic similarity, and merge multiple target description texts into a final target description text according to the weights.

[0175] As one possible implementation, the descriptive text correction system 400 can also output and display multiple target descriptive texts, receive user selection instructions, and determine the target descriptive text to be fused. The target descriptive text to be fused is the portion of the target descriptive text selected by the user from multiple target descriptive texts. The descriptive text correction system 400 can perform semantic fusion based on the target descriptive text to be fused to obtain the fused descriptive text.

[0176] Based on the above Figure 7 The descriptions of steps 701 to 703 describe how the text correction system 400, after acquiring multiple description texts generated by multiple VLMs, corrects the initial description texts based on semantic differences. By considering these semantic differences, the corrected target description texts avoid discrepancies and descriptive biases, improving the accuracy of the video descriptions. Furthermore, text semantic fusion is performed on the multiple target description texts to obtain the fused description text, further improving the quality of the final generated fused description text.

[0177] Furthermore, since the fused description text generated by the description text correction system 400 is obtained by fusing multiple target description texts after correcting multiple initial description texts generated by multiple VLMs, the quality of the generated fused description text no longer depends on the upper limit of the capability of a single model, thus avoiding the problem of computational resource limitations that exist when a single model generates the final description text.

[0178] The above describes the overall process of describing text correction methods. Next, we will combine... Figures 10 to 12 The process of each VLM correcting errors in the description text is explained in detail.

[0179] The first difference result includes at least one difference item, and each difference item in the at least one difference item is used to indicate a semantic difference content of a type in a hierarchy. Different difference items can be of the same type belonging to the same hierarchy, or different types belonging to the same hierarchy, or the same type belonging to different hierarchy, or different types belonging to different hierarchy. The embodiments of this application do not limit this.

[0180] The descriptive text correction system 400 can generate a first correction prompt word corresponding to each difference item. The first correction prompt word corresponding to each difference item is used to guide the first VLM to correct each difference item in the first descriptive text.

[0181] Specifically, the descriptive text correction system 400 can generate a first correction prompt word for each difference item based on the prompt word template corresponding to the level and type to which each difference item belongs. Alternatively, the descriptive text correction system 400 can generate a first correction prompt word for each difference item by inputting each difference item into the prompt word generation model.

[0182] The process by which the text correction system 400 generates the first correction prompt word corresponding to each difference item can be found in the above text. Figure 7 Step 703 and below Figures 10 to 12 The relevant descriptions in the original text will not be repeated in the embodiments of this application.

[0183] As one possible implementation, the descriptive text correction system 400 can create a dialogue corresponding to each difference item in the first VLM for each difference item correction. The descriptive text correction system 400 can input a video, first descriptive text, each difference item, and a corresponding first correction prompt word into the dialogue corresponding to each difference item in the first VLM. Based on the first VLM and according to the first correction prompt word and video for each difference item, the descriptive text correction system 400 corrects each difference item in the first descriptive text, obtaining the first intermediate descriptive text after correction in this round of correction iteration output by the first VLM.

[0184] Specifically, the text correction system 400 generates a correction task corresponding to each difference item in the first difference result, and forms a correction task group corresponding to the first difference result. The context of each correction task corresponding to a difference item includes the difference item, the first correction prompt word corresponding to the difference item, and the video.

[0185] The text correction system 400 can also output correction tasks, which users can modify, supplement, and manually confirm to obtain updated correction tasks, thereby improving the accuracy of the correction tasks.

[0186] The descriptive text correction system 400 can create a separate dialogue for each correction task in the first Virtual Library (VLM) for each difference item in the correction task group. The descriptive text correction system 400 executes each correction task, inputting the difference item, the first correction prompt word corresponding to the difference item, and the video in the context of each correction task into the dialogue corresponding to each correction task. Based on the first VLM and the input context, the descriptive text correction system 400 performs correction, obtaining the first intermediate descriptive text output by the first VLM after correcting at least one difference item in the first descriptive text in this round of correction iteration.

[0187] For example, when the difference item belongs to the replacement difference type at the global level, the semantics of the first error correction prompt word generated by the description text correction system 400 can be "the first VLM determines whether the video content occurs in environment A, rather than environment B described in the second description text".

[0188] In the case where the difference item belongs to the omission difference type at the global level, the semantics of the first error correction prompt word generated by the description text correction system 400 can be "The first VLM determines whether the occurrence background of the video content described in the second description text is environment B".

[0189] When the difference item belongs to the newly added difference type at the global level, the semantics of the first error correction prompt word generated by the description text correction system 400 can be "the first VLM determines whether the background of the occurrence of video content omitted by the second description text is environment A".

[0190] Figure 10 This is a schematic diagram of a generated error correction task group provided in an embodiment of this application.

[0191] like Figure 10 As shown, VLM1 corresponds to error correction task group 1 and error correction task group 2 to be executed.

[0192] The descriptive text correction system 400 generates a correction task consisting of each difference item in the difference results between descriptive text 1 and descriptive text 2, along with the corresponding correction prompt word and video. The correction task corresponding to each difference item forms a correction task group 1 to be executed.

[0193] The descriptive text correction system 400 generates a correction task consisting of each difference item in the difference results between descriptive text 1 and descriptive text 3, along with the corresponding correction prompt word and video. The correction tasks corresponding to each difference item form a group of correction tasks to be executed, 2.

[0194] VLM2 corresponds to error correction task group 3 and error correction task group 4 that are yet to be executed.

[0195] The descriptive text correction system 400 generates a correction task consisting of each difference item in the difference results between descriptive text 2 and descriptive text 1, along with the corresponding correction prompt word and video. The correction tasks corresponding to each difference item form a group of correction tasks to be executed, 3.

[0196] The descriptive text correction system 400 generates a correction task consisting of each difference item in the difference results between descriptive text 2 and descriptive text 3, along with the corresponding correction prompt word and video. The correction tasks corresponding to each difference item form a group of correction tasks 4 to be executed.

[0197] VLM3 corresponds to error correction task group 5 and error correction task group 6 that are yet to be executed.

[0198] The descriptive text correction system 400 generates a correction task consisting of each difference item in the difference results between descriptive text 3 and descriptive text 1, along with the corresponding correction prompt word and video. The correction task corresponding to each difference item forms a group of correction tasks 5 to be executed.

[0199] The descriptive text correction system 400 generates a correction task consisting of each difference item in the difference results between descriptive text 3 and descriptive text 2, along with the corresponding correction prompt word and video. The correction task corresponding to each difference item forms a group of correction tasks 6 to be executed.

[0200] The error correction tasks in the error correction task group can be found in Table 4 below, and will not be elaborated upon here.

[0201] Figure 11 This is a schematic diagram of an error correction task group provided in an embodiment of this application.

[0202] like Figure 11 As shown, the description text correction system 400 executes correction task group 1 and correction task group 2, inputs the correction tasks in correction task group 1 and correction task group 2 into VLM1, and obtains the corrected intermediate description text 1 output by VLM1.

[0203] The description text correction system 400 executes correction task group 3 and correction task group 4, inputs the correction tasks in correction task group 3 and correction task group 4 into VLM2, and obtains the corrected intermediate description text 2 output by VLM2.

[0204] The description text correction system 400 executes correction task groups 5 and 6, inputs the correction tasks in correction task groups 5 and 6 into VLM3, and obtains the corrected intermediate description text 3 output by VLM3.

[0205] As one possible implementation, the error correction process of the text correction system 400 varies depending on the level to which the difference item belongs. When the difference item belongs to the global level, the steps of the text correction system 400 in performing the error correction can be referred to the relevant descriptions in the possible implementations above. When the difference item belongs to the local level, the relevant descriptions in the first and second optional implementations below can be referred to.

[0206] Among them, the local level refers to other levels besides the global level, that is, the local level can include at least one of the object level, scene level, target level or action level.

[0207] As a first optional implementation, when the difference item belongs to the local level, the descriptive text correction system 400 can determine the time range corresponding to the target difference item at the local level in the video. The descriptive text correction system 400 can input the video, time range, first descriptive text, target difference item, and corresponding first correction prompt word in the dialogue corresponding to the target difference item. Based on the first VLM, the descriptive text correction system 400 can correct the target difference item in the first descriptive text according to the first correction prompt word corresponding to the target difference item and the video segment corresponding to the time range in the video.

[0208] Specifically, the text correction system 400 can determine the time range of corresponding target differences in the video based on a first VLM or a second VLM. The time range can be represented in the format of [start time, end time], and the start and end times can be in seconds expressed as floating-point numbers; this embodiment does not impose any limitations on this.

[0209] When the target difference item belongs to the local level of replacement difference type, the description text correction system 400 obtains the time range corresponding to the target difference item determined by the first VLM.

[0210] When the target difference item belongs to the omission difference type at the local level, the description text correction system 400 obtains the time range corresponding to the target difference item determined by the second VLM that generates the second description text.

[0211] When the target difference item belongs to a newly added difference type at the local level, the descriptive text correction system 400 obtains the time range corresponding to the target difference item determined by the first VLM. Additionally, the descriptive text correction system 400 can also provide this time range to the second VLM as the time range for the second VLM to determine the target difference item belonging to the omitted difference type.

[0212] The descriptive text correction system 400 generates a corresponding target correction task based on the target discrepancies. The context of the target correction task may include the target discrepancies, corresponding prompts, video, time range, and a first descriptive text. The descriptive text correction system 400 creates a dialog for the target correction task in a first Virtual Library (VLM) and executes the target correction task, inputting the target discrepancies, corresponding prompts, video, and time range from the context of the target correction task into the dialog. Based on the first VLM, the descriptive text correction system 400 corrects the target discrepancies in the first descriptive text according to the input context of the target correction task.

[0213] Based on the above optional implementation methods, the description text correction system 400 inputs the time range into the first VLM, and the first VLM reviews the video segments within the time range, realizing the temporal scaling from panoramic to local, simulating the logic of human processing complex information, thereby improving the quality of the first intermediate description text generated after error correction.

[0214] As a second optional implementation, when the difference item belongs to the target level, the descriptive text correction system 400 can determine the time range corresponding to the target difference item at the target level in the video. The descriptive text correction system 400 can then crop video segments corresponding to the time range of the target difference item from the video. In the dialogue corresponding to the target difference item, the descriptive text correction system 400 inputs the video segment, the first descriptive text, the target difference item, and the corresponding first correction prompt. Based on the first VLM (Visual Model), the descriptive text correction system 400 corrects the target difference item in the first descriptive text according to the first correction prompt and the video segment.

[0215] Specifically, the descriptive text correction system 400 segments the video to obtain video segments corresponding to a time range. Based on the target difference item, the descriptive text correction system 400 generates a corresponding target correction task. The context of the target correction task may include the target difference item, the first correction prompt word corresponding to the target difference item, the video segment, and the first descriptive text. The descriptive text correction system 400 creates a dialogue for the target correction task in the first Video Library (VLM) and executes the target correction task. Based on the target sampling frame number per unit duration, the descriptive text correction system 400 extracts frames from the video segment using the first VLM to obtain target video frames. The target sampling frame number is greater than the sampling frame number per unit duration extracted by the first VLM in the video. Based on the first VLM, the descriptive text correction system 400 corrects the target difference item in the first descriptive text according to the first correction prompt word corresponding to the target difference item and the target video frame.

[0216] For example, the target difference item of an object type corresponds to a 3-second video segment in the time range of [2.35, 5.35], where the video is 16 frames per second, and the 3-second video segment contains a total of 48 frames. During the video frame extraction and sampling stage of the VLM, in order to meet the input context length requirement of the first VLM, the number of sampling frames per unit duration in the video is 2.

[0217] However, during the first VLM's correction of the target difference item, the descriptive text correction system 400 only performs correction processing on this 3-second video segment. If the sampling frame number of 2 per unit duration in the video is continued, that is, only 2 frames are extracted per second for the 3-second video segment, the first VLM only obtains 6 frames for judging the content of the video segment, which is insufficient to meet the video frame data required for correcting the target difference item. Therefore, the target sampling frame number per unit duration is 8, which is an increase compared to the sampling frame number per unit duration when extracting frames from the entire video. The descriptive text correction system 400 extracts a total of 24 frames in the 3-second video segment corresponding to the target difference item, and corrects the target difference item of the first descriptive text based on the first VLM according to the 24 frames, obtaining the corrected first intermediate descriptive text output.

[0218] By extracting more video frames from the 3-second video clip corresponding to the target difference item, the video clip can be more precisely identified, thus ensuring the accuracy of the first intermediate description text after error correction.

[0219] Figure 12 This is a schematic diagram illustrating a refined frame extraction of a video segment corresponding to a target difference item, provided as an embodiment of this application.

[0220] like Figure 12 As shown, the overall video includes object-level replacement differences, missing differences, and added differences. When the first VLM extracts video frames from the overall video, it extracts 3 frames from the video segment with replacement differences, 4 frames from the video segment with missing differences, and 3 frames from the video segment with added differences.

[0221] The text correction system 400 describes that after extracting video segments from video frames, the first VLM extracts refined video frames from the video segments. It extracts 7 frames from video segments with replaced differences, 7 frames from video segments with missing differences, and 7 frames from video segments with newly added differences.

[0222] In this way, the number of refined video frames extracted from the video segment corresponding to the target difference item increases compared to the number of video frames extracted from the whole video. This allows for error correction of the target difference item of the first descriptive text based on more video frames, ensuring the accuracy of the first intermediate descriptive text after error correction.

[0223] As another possible implementation, the descriptive text correction system 400 can determine the time range corresponding to the target difference item at the target level in the video based on the level to which at least one difference item in the first difference result belongs. The descriptive text correction system 400 inputs the video, the first descriptive text, at least one difference item, the corresponding first correction prompt word, and the time range corresponding to the target difference item into a one-round dialogue of the first VLM for correction, and obtains the corrected first intermediate descriptive text output by the first VLM.

[0224] Specifically, the text correction system 400 generates a correction task corresponding to each difference item in the first difference result, and forms a correction task group corresponding to the first difference result. The correction task corresponding to the difference item at the global level includes the difference item, the first correction prompt word corresponding to the difference item, and the video. The target correction task corresponding to the target difference item at the target level includes the target difference item, the first correction prompt word corresponding to the target difference item, and the time range.

[0225] The description text correction system 400 can create a dialogue for the correction task group, input the context of all correction tasks in the correction task group into a round of dialogue of the first VLM for batch correction, and obtain the first intermediate description text after correction output by the first VLM.

[0226] The foregoing described the detailed process by which multiple VLMs correct the initial description text they generate. The following will provide a detailed explanation of the description text correction process of this application embodiment through specific examples.

[0227] In a video depicting a robotic arm operating a vegetable shelf, the descriptive text correction system 400 generates multiple descriptive texts for the video using multiple VLMs (Visual Models). A single frame from the video, such as... Figure 13 As shown, Figure 13 This is a schematic diagram of a video frame showing a robot arm operating a vegetable shelf, as provided in an embodiment of this application.

[0228] Among them, VLM1 is the description text 1 generated by the video, which is "The scene is set on a vegetable shelf. A plastic bag is tied to a supermarket shopping cart. The robotic arm on the right picks up a green mushroom from the shelf and puts it in the plastic bag, then picks up a yellow corn and puts it in the plastic bag, and finally picks up a purple vegetable and puts it in the plastic bag."

[0229] The description text 2 generated by VLM2 for the video is "A white robotic arm with black grippers picked up white mushrooms, yellow corn and purple cabbage and put them in a plastic bag."

[0230] The description text 3 generated by VLM3 for the video is: "In the vegetable section of a supermarket, the right robotic arm first picks up yellow corn and puts it into a plastic bag on the supermarket shopping cart, then picks up purple onions and puts them into the plastic bag as well. During the process, the left robotic arm does not move."

[0231] The descriptive text correction system 400 inputs descriptive text 2 into VLM1, performs semantic difference analysis on descriptive text 1 and descriptive text 2 based on VLM1, and obtains a difference result table between descriptive text 1 and descriptive text 2. The descriptive text correction system 400 inputs descriptive text 3 into VLM1, performs semantic difference analysis on descriptive text 1 and descriptive text 3 based on VLM1, and obtains a difference result table between descriptive text 1 and descriptive text 3.

[0232] Table 2 shows the differences between description text 1 and description text 2 provided in the embodiments of this application.

[0233] Table 2

[0234] As shown in Table 2, among the differences between description text 1 and description text 2, the difference item at the object level for the replacement difference type is "green mushroom / white mushroom," meaning description text 1 describes green mushrooms, while description text 2 describes white mushrooms. The difference item at the global level for the new difference type is "vegetable shelf scene," meaning description text 1 describes a vegetable shelf scene, but description text 2 does not describe a vegetable shelf scene.

[0235] Table 3 shows the differences between description text 1 and description text 3 provided in the embodiments of this application.

[0236] Table 3

[0237] As shown in Table 3, among the differences between description text 1 and description text 3, the content of the object-level replacement difference type difference item is "purple vegetables / purple onions," meaning that description text 1 describes purple vegetables, while description text 3 describes purple onions. The content of the object-level new difference type difference item is "green mushrooms," meaning that description text 1 describes green mushrooms, but description text 3 does not describe green mushrooms.

[0238] The text correction system 400 determines the time range of the video corresponding to the target difference item at the object level in the difference result table. The text correction system 400 generates the time range of each difference item in the first difference result and assembles it into an information table.

[0239] Table 4 is an information table of time ranges for differences provided in an embodiment of this application.

[0240] Table 4

[0241] As shown in Table 4, the time range corresponding to the first difference item is [0, 48], the content is "vegetable shelf scene", the level is global level, and the type is new difference type.

[0242] The time range corresponding to Difference Item 2 is [0, 14], the content is "green mushroom / white mushroom", the level is the object level, and the type is replacement difference type.

[0243] The time range corresponding to Difference Item 3 is [28, 38], the content is "purple vegetables / purple onions", the level is the object level, and the type is replacement difference type.

[0244] The time range corresponding to difference item 4 is [0, 14], the content is "green mushroom", the level is the object level, and the type is new difference type.

[0245] The text correction system 400 generates a group of correction tasks based on each difference item in Table 4.

[0246] Specifically, the text correction system 400 generates four error-correction tasks corresponding to the four differences in VLM1 based on the information table shown in Table 4.

[0247] In the error correction task generated for difference item 1, the first step is for the description text error correction system 400 to submit the video, description text 1, difference item 1, and the corresponding first error correction prompt to VLM1. The first error correction prompt corresponding to difference item 1 requires VLM1 to perform an accuracy check on the "vegetable shelf scene" to determine if it exists.

[0248] In the error correction task generated for difference item 2, the first step is that the description text error correction system 400 extracts video segment 1 with a time range of [0, 14] from the video based on the time range of difference item 2 [0, 14]. The second step is that the description text error correction system 400 submits video segment 1, description text 1, difference item 2, and the corresponding first error correction prompt to VLM1. The first error correction prompt corresponding to difference item 2 requires VLM1 to perform an accuracy check on "green mushroom" to determine whether it is "white mushroom".

[0249] In the error correction task generated for difference item 3, the first step is that the description text error correction system 400 extracts video segment 2 with a time range of [28, 38] from the video based on the time range of difference item 3 [28, 38]. The second step is that the description text error correction system 400 submits video segment 2, description text 1, difference item 3, and the corresponding first error correction prompt to VLM1. The first error correction prompt corresponding to difference item 3 requires VLM1 to perform an accuracy check on "purple vegetables" to determine whether it is "purple onions".

[0250] In the error correction task generated for difference item 4, the first step is that the description text error correction system 400 extracts video segment 3 with a time range of [0, 14] from the video based on the time range of difference item 4 [0, 14]. The second step is that the description text error correction system 400 submits video segment 3, description text 1, difference item 4, and the corresponding first error correction prompt to VLM1. The first error correction prompt corresponding to difference item 4 requires VLM1 to perform an accuracy check on "green mushroom" to determine whether it exists.

[0251] The description text correction system 400 executes the correction task group and obtains the corrected intermediate description text 1 output by VLM1.

[0252] The description text correction system 400 performs the above steps on VLM2 and VLM3 respectively, and obtains the corrected intermediate description text 2 output by VLM2 and the corrected intermediate description text 3 output by VLM3.

[0253] The descriptive text correction system 400, after at least one round of error correction iteration, obtains three corrected target descriptive texts from the three VLMs output in the last round of error correction iteration when the error correction iteration stopping condition is met. Based on user instructions, if the instruction is a fusion instruction, the descriptive text correction system 400 fuses the three corrected target descriptive texts to obtain the final fused descriptive text.

[0254] The preceding text described the text correction process in detail using a specific example. The following section will combine... Figure 5 The text correction system 400 describes the modules and the execution flow of each module in the text correction method.

[0255] The execution flow of the text exchange and analysis module 420 is described below: Input: Input information The caption text output by the large model itself: VLM_self_caption; Other large model output caption text arrays: VLM_oth_captions: {VLM_0_caption,VLM_2_caption, ... , VLM_n_caption}.

[0256] Step 1: Initialize caption analysis prompt: prompt = "This is your output video description: " + ${VLM_self_caption} Step 2: Connect the caption content of other models if (length(VLM_oth_captions)>0) { N = length(VLM_oth_captions) prompt += "There are other " + N + " models that helped me output video descriptions: " for (int i = 0; i <N; i++) { prompt += "The video description given by the " + i + " model is as follows: " + VLM_oth_captions[i] + ”."; } } Step 3: Output the caption difference table according to the required format: prompt += "Please compare your video description with the video descriptions given by other models, and summarize the semantic differences between your description and those of other models. Fill these differences into N tables with 3 columns and 2 rows. The columns are the differences in the following 3 categories: replacement difference (different descriptions of the same object or scene), omission difference (objects or scenes described by others but not by yourself), or addition difference (objects or scenes described by yourself but not by others). The rows are to determine whether these differences belong to the global scope or the object scope." The text exchange and analysis module 420 performs the above process for each VLM to obtain a difference result table. The difference result table can be referenced from the examples in Tables 2 and 3 above; it will not be elaborated further in this embodiment.

[0257] The execution process of the error correction task generation module 430 is as follows: Input: Input information Total video duration T Replace the difference array sub_diffs of the difference type: {sub_diff_1(object), sub_diff_2(scene), ..., sub_diff_(scene)} The difference array del_diffs, which is missing the difference type, is: {del_diff_1(object), del_diff_2(scene), ..., del_diff_p(scene)} A new difference array ins_diffs for the difference type is added: {ins_diff_1(object), ins_diff_2(scene), ..., ins_diff_p(scene)} Step 1: Initialization prompt: prompt = "You are now required to analyze the semantic differences in the following video description:" Step 2: Concatenate the task description if (length(sub_diffs)>0) { O = length(sub_diffs) Prompt += "The total number of semantic differences of the replacement difference type is: " + O + ". for (int i = O; i <O; i++) { action = "The " + i + "th replacement class difference:" if (sub_diffs[i].label == "object") { action = "You need to provide video content" + sub_diffs[i] + "'s start and end times, returned in the format: {start_time:12, end_time: 16}, in seconds." } else { action = "This difference requires viewing the full video transcript; you can return {start_time:0, end_time:T}" } prompt += action; } } if (length(del_diffs)>0) { P = length(del_diffs) Prompt += "The total number of semantic differences of the omitted difference type is: " + P + ". for (int i = P; i <P; i++) { action = "The " + i + "th omitted class difference:" if (sub_diffs[i].label == "object") { action = “Your video content description is missing an operation” + sub_diffs[i] + “,Please confirm if it actually exists, and provide the start and end times of the video content. Return the result in the format: {start_time:12, end_time:16} in seconds.” } else { action = "This difference requires viewing the full video; you can return {start_time:0, end_time:T}" } prompt += action; } } if (length(ins_diffs)>0) { Q = length(ins_diffs) Prompt += "The total number of semantic differences for the newly added difference types is: " + Q + ". for (int i = Q; i <Q; i++) { action = "The " + i + "th new class difference:" if (sub_diffs[i].label == "object") { action = "An operation has been added to your video content description" + sub_diffs[i] + ", Please confirm if it actually exists, and provide the start and end times of the video content. The returned value will be in the format {start_time:12, end_time:16}, in seconds." } else { action = "This difference requires viewing the full video; you can return {start_time:0, end_time:T}" } prompt += action; } } Step 3: Return the execution task table as required: prompt += "Please return the specific time of the object operation according to the above format requirements. If the fact check did not occur, the time result does not need to be returned. The model needs to return a table with {O+P+Q} rows, each row being a difference item. The table has 5 columns, namely start time, end time, difference content, difference category (replaced difference, missing difference, added difference), and difference level (global and object)." Specifically, the difference array for replacing difference types can be composed of difference items of the same type from the difference results table. The difference array for missing difference types can be composed of difference items of the same type from the difference results table. The difference array for newly added difference types can be composed of difference items of the same type from the difference results table.

[0258] The error correction task generation module 430 generates an information table consisting of the time range of the difference items in the difference results determined by each VLM. The information table can be referenced from the example in Table 4 above; this embodiment will not be elaborated upon here.

[0259] The execution process of the error correction task module 440 is as follows: Input: Input information The array of tasks to be executed is named tasks: {task1, task2, ..., taskn}. The video segments corresponding to the task array to be executed: task_videos:{task_video_1, task_video_2,…, task_video_n}} Step 1: Initialize and execute the task prompt: prompt = "You need to update your video description according to the following tasks." Step 2: Concatenate the description of the error correction task. if (length(tasks)>0) { N = length(tasks) for (int i = 0; i <N; i++) { prompt += "The " + i + "task is:" if (tasks[i].diff_class == "Replace diff type") { prompt += “Please carefully check the “ + i + “ video segment and perform an accuracy check on the video content “ + VLM_oth_captions[i] + “.; } else if (tasks[i].diff_class == "missing difference type") { prompt += "Please carefully check the " + i + " video segment and determine whether any content is missing from the video content " + VLM_oth_captions[i] + ".; } else { prompt += "Please carefully check the " + i + " video segment and determine if the video content " + VLM_oth_captions[i] + " exists."; } } } Step 3: Require the model to perform a Recaption: prompt += "Please carefully review the video content again based on the above task and the re-uploaded video clip, and update the video description." The array of tasks to be executed can be a group of error correction tasks generated by the text correction system 400 based on the differences in the output of each VLM. Examples of error correction task groups can be found in Table 4 above, and will not be elaborated upon here.

[0260] The error correction task module 540 executes each error correction task group to obtain the error-corrected intermediate description text output by each VLM.

[0261] The description text error correction method provided in the embodiments of this application has been described in detail above. The following section, in conjunction with... Figure 14 This application describes a text correction device 1400 provided in an embodiment of the present application. The text correction device 1400 is used to implement the above-described... Figures 4 to 13 The implementation shown describes the functionality of the text correction system 400.

[0262] Figure 14 This is a schematic diagram illustrating a text correction device 1400 provided in an embodiment of this application. Figure 14 As shown, the text correction device 1400 includes an acquisition module 1401, a correction module 1402, and a fusion module 1403.

[0263] The acquisition module 1401 is used to acquire multiple initial description texts generated by multiple VLMs for the video. For example, the acquisition module 1401 is used to execute the above... Figure 7 Step 701 in the process.

[0264] Error correction module 1402 is used to correct multiple initial description texts based on semantic differences between them, resulting in multiple corrected target description texts. For example, error correction module 1402 is used to execute the above... Figure 7 Step 702 in the process.

[0265] The fusion module 1403 is used to perform text semantic fusion on multiple target description texts to obtain the fused description text. For example, the fusion module 1403 is used to execute the above... Figure 7 Step 703 in the process.

[0266] As one possible implementation, the error correction module 1402 is specifically used to perform at least one round of error correction iteration on multiple initial description texts based on the semantic differences between multiple initial description texts, so as to obtain multiple target description texts after error correction.

[0267] In the first round of error correction iteration, based on the semantic differences between multiple initial description texts, multiple initial description texts are corrected to obtain multiple intermediate description texts for the first round of error correction iteration. In the m-th round of error correction iteration, based on the semantic differences between multiple intermediate description texts obtained in the (m-1)-th round of error correction iteration, multiple intermediate description texts obtained in the (m-1)-th round of error correction iteration are corrected to obtain multiple intermediate description texts for the m-th round of error correction iteration, where m is a positive integer greater than or equal to 2. The multiple target description texts are the multiple intermediate description texts obtained in the last round of error correction iteration when the error correction iteration stopping condition is met.

[0268] As one possible implementation, the stopping condition for error correction iteration is: the semantic difference rate of the multiple intermediate description texts obtained in the current round of error correction iteration is less than the semantic difference rate threshold; or, the number of error correction iteration rounds reaches the round number threshold.

[0269] As one possible implementation, the error correction module 1402 is specifically used to perform text semantic difference analysis based on multiple descriptive texts in any round of error correction iteration, to obtain the difference result corresponding to the current round of error correction iteration, and the difference result is used to indicate the semantic difference between every two descriptive texts in the multiple descriptive texts; and to perform error correction based on the difference result corresponding to the current round of error correction iteration and multiple descriptive texts, to obtain multiple intermediate descriptive texts for the current round of error correction iteration.

[0270] In the first round of error correction iteration, the multiple description texts are multiple initial description texts; in the m-th round of error correction iteration, the multiple description texts are multiple intermediate description texts obtained in the (m-1)-th round of error correction iteration.

[0271] As one possible implementation, the error correction module 1402 is specifically used to input multiple descriptive texts into the difference analysis model to obtain the difference results of every two descriptive texts in the multiple descriptive texts corresponding to the current error correction iteration.

[0272] As one possible implementation, the error correction module 1402 is specifically used to input the first description text and the second description text into the first VLM for text semantic difference analysis, and obtain the first difference result corresponding to the current error correction iteration output by the first VLM. The first difference result is used to indicate the semantic difference between the first description text and the second description text. The difference result includes the first difference result. In the first round of error correction iteration, the first description text is the initial description text generated by the first VLM for the video, and the second description text is the initial description text generated by the second VLM for the video; in the m-th round of error correction iteration, the first description text and the second description text are the intermediate description texts obtained from the error correction in the (m-1)-th round of error correction iteration. Based on the first VLM, the first description text is corrected according to the first difference result, and the first intermediate description text corresponding to this round of error correction iteration is obtained.

[0273] As one possible implementation, the error correction module 1402 is specifically used to generate a first error correction prompt word based on the first difference result. The first error correction prompt word is used to guide the first VLM to correct the first description text. The video, the first description text and the first error correction prompt word are input into the first VLM. The first VLM corrects the first description text to obtain the first intermediate description text after correction in this round of error correction iteration output by the first VLM.

[0274] As one possible implementation, the first difference result is used to indicate different levels and types of semantic differences between the first descriptive text and the second descriptive text. The levels include at least one of the video's global, object, scene, target, or action, and the types include at least one of the replacement difference, omission difference, or addition difference. Replacement differences are used to indicate semantic differences that are different between the first description text and the second description text. Omission differences are used to indicate semantic differences that are not described in the first description text but are described in the second description text. Added differences are used to indicate semantic differences that are described in the first description text but not in the second description text.

[0275] As one possible implementation, the first difference result includes at least one difference item, each of the at least one difference item indicating a semantic difference content of a type in a hierarchy. The error correction module 1402 is specifically used to generate a first error correction prompt word corresponding to each difference item.

[0276] As one possible implementation, the error correction module 1402 is specifically used to input a video, a first descriptive text, each difference item, and a corresponding first error correction prompt word in the dialogue corresponding to each difference item in the first VLM; based on the first VLM, the first VLM corrects at least one difference item of the first descriptive text according to the video and the first error correction prompt word corresponding to at least one difference item, and obtains the first intermediate descriptive text output by the first VLM in this round of error correction iteration.

[0277] As one possible implementation, the error correction module 1402 is specifically used to crop video segments corresponding to the time range of the target difference item in the video according to the time range of the local level in the video, where the local level is other levels besides the global level; in the dialogue corresponding to the target difference item in the first VLM, the video segment, the first descriptive text, the target difference item and the corresponding first error correction prompt word are input.

[0278] As one possible implementation, the error correction module 1402 is specifically used to correct the target difference item of the first descriptive text based on the first VLM according to the first error correction prompt word and video segment corresponding to the target difference item.

[0279] As one possible implementation, the error correction module 1402 is specifically used to extract frames from the video segment based on the first VLM according to the target number of sampling frames per unit duration to obtain the target video frame; the target number of sampling frames is greater than the number of sampling frames per unit duration extracted by the first VLM in the video; based on the first VLM, the target difference item of the first descriptive text is corrected according to the first error correction prompt word corresponding to the target difference item and the target video frame.

[0280] As one possible implementation, the error correction module 1402 is specifically used to input the video, the first descriptive text, at least one difference item and the corresponding first error correction prompt word, and the time range corresponding to the target difference item at the local level in the video into a round of dialogue of the first VLM, where the local level refers to other levels besides the global level.

[0281] As one possible implementation, the fusion module 1403 is specifically used to input multiple target description texts into a large language model (LLM), perform semantic fusion based on the large language model, and obtain the fused description text.

[0282] As one possible implementation, the fusion module 1403 is specifically used to perform semantic fusion of multiple target description texts using a weighted fusion algorithm to obtain the fused description text.

[0283] It should be understood that the text correction device 1400 described in this application embodiment can be implemented by a graphics processing unit (GPU), a neural network processing unit (NPU), an application-specific integrated circuit (ASIC), or a programmable logic device (PLD). The PLD can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. It is implemented through software. Figures 4 to 13 When describing the function of the text correction system 400 in the implementation shown, the text correction device 1400 and its various modules can also be described as software modules.

[0284] It should be understood that the acquisition module 1401, error correction module 1402, and fusion module 1403 in the text correction device 1400 described in this application embodiment can also be implemented in software or in hardware. For example, the implementation of the acquisition module 1401 will be described below. Similarly, the implementation of the error correction module 1402 and the fusion module 1403 can refer to the implementation of the acquisition module 1401.

[0285] As an example of a software functional unit, module 1401 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, module 1401 may include code running on multiple hosts / virtual machines / containers. It should be understood that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0286] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0287] As an example of a hardware functional unit, the acquisition module 1401 may include at least one computing device, such as a server. Alternatively, the acquisition module 1401 may be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0288] The multiple computing devices included in the acquisition module 1401 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the acquisition module 1401 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the acquisition module 1401 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0289] It should be understood that in other embodiments, the acquisition module 1401 can be used to execute any step in the descriptive text correction method, and the correction module 1402 and the fusion module 1403 can be used to execute any step in the descriptive text correction method. The steps implemented by the acquisition module 1401, the correction module 1402 and the fusion module 1403 can be specified as needed. By implementing different steps in the descriptive text correction method through the acquisition module 1401, the correction module 1402 and the fusion module 1403 respectively, all functions of the descriptive text correction device 1400 can be realized.

[0290] This application also provides an electronic device, please refer to... Figure 15 , Figure 15This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 1500 includes a memory 1501, a processor 1502, a communication interface 1503, and a bus 1504. The memory 1501, the processor 1502, and the communication interface 1503 are interconnected via the bus 1504.

[0291] Memory 1501 may be a read-only memory, a static storage device, a dynamic storage device, or a random access memory. Memory 1501 may store computer instructions. When the computer instructions stored in memory 1501 are executed by processor 1502, processor 1502 and communication interface 1503 are used to perform the steps describing the text correction method. For example, processor 1502 is used to perform the above... Figures 4 to 13 The text correction system 400 is described in the text description, and the above Figure 14 The diagram illustrates the function of the text correction device 1400.

[0292] Processor 1502 may be a general-purpose CPU, an application-specific integrated circuit (ASIC), a GPU, or any combination thereof. Processor 1502 may include one or more chips.

[0293] The communication interface 1503 uses a transceiver module, such as, but not limited to, a transceiver, to enable communication between the electronic device 1500 and other devices or communication networks.

[0294] Bus 1504 may include a pathway for transmitting information between various components of electronic device 1500 (e.g., memory 1501, processor 1502, communication interface 1503).

[0295] Electronic device 1500 can be a computer (e.g., a server) in a cloud data center, or a computer in an edge data center, or a terminal. For example, electronic device 1500 can be a mobile phone, tablet computer, etc.

[0296] This application also provides a computing device 1600. Figure 16 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application, such as... Figure 16 As shown, the computing device 1600 includes a bus 1602, a processor 1604, a memory 1606, and a communication interface 1608. The processor 1604, the memory 1606, and the communication interface 1608 communicate with each other via the bus 1602. The computing device 1600 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1600.

[0297] The 1602 bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 16 The bus 1602 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 1602 may include a path for transmitting information between various components of the computing device 1600 (e.g., memory 1606, processor 1604, communication interface 1608).

[0298] Processor 1604 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0299] The memory 1606 may include volatile memory, such as random access memory (RAM). The processor 1604 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0300] The memory 1606 stores executable program code, and the processor 1604 executes this executable program code to implement the functions of the aforementioned acquisition module 1401, error correction module 1402, and fusion module 1403, thereby realizing the text correction method. That is, the memory 1606 stores instructions for executing the text correction method.

[0301] The communication interface 1608 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 1600 and other devices or communication networks.

[0302] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0303] Figure 17 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application, such as... Figure 17 As shown, the computing device cluster includes at least one computing device 1600. The memory 1606 of one or more computing devices 1600 in the computing device cluster may store the same instructions for executing methods describing text correction.

[0304] In some possible implementations, the memory 1606 of one or more computing devices 1600 in the computing device cluster may also store partial instructions for executing the text correction method. In other words, a combination of one or more computing devices 1600 can jointly execute the instructions for executing the text correction method.

[0305] It should be understood that the memory 1606 in different computing devices 1600 within the computing device cluster can store different instructions, each used to execute a portion of the functions describing the text correction device 1400. That is, the instructions stored in the memory 1606 of different computing devices 1600 can implement the functions of one or more modules among the acquisition module 1401, the error correction module 1402, and the fusion module 1403.

[0306] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 18 One possible implementation method is shown. Figure 18 This application provides a schematic diagram of a network connection structure between computing devices, as shown in the embodiments of the present application. Figure 18 As shown, the two computing devices 1600A and 1600B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this possible implementation, the memory 1606 in computing device 1600A stores instructions for executing the functions of the acquisition module 1401. Simultaneously, the memory 1606 in computing device 1600B stores instructions for executing the functions of the error correction module 1402 and the fusion module 1403.

[0307] It should be understood that Figure 18The functions of computing device 1600A shown can also be performed by multiple computing devices 1600. Similarly, the functions of computing device 1600B can also be performed by multiple computing devices 1600.

[0308] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform a descriptive text error correction method.

[0309] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to perform a descriptive text error correction method.

Claims

1. A method for describing text error correction, characterized in that, The method includes: Obtain multiple initial descriptive texts generated by multiple Visual Language Models (VLMs) for the video; Based on the semantic differences between the multiple initial description texts, the multiple initial description texts are corrected to obtain multiple target description texts after correction. The multiple target description texts are semantically fused to obtain the fused description text.

2. The method according to claim 1, characterized in that, The step of correcting the semantic differences between the multiple initial description texts to obtain multiple corrected target description texts includes: Based on the semantic differences between the multiple initial description texts, at least one round of error correction iteration is performed on the multiple initial description texts to obtain the multiple target description texts after error correction. In the first round of error correction iteration, based on the semantic differences between the multiple initial description texts, the multiple initial description texts are corrected to obtain multiple intermediate description texts for the first round of error correction iteration; During the m-th round of error correction iteration, based on the semantic differences between the multiple intermediate description texts obtained in the (m-1)-th round of error correction iteration, the multiple intermediate description texts obtained in the (m-1)-th round of error correction iteration are corrected to obtain the multiple intermediate description texts in the m-th round of error correction iteration, where m is a positive integer greater than or equal to 2. The multiple target description texts are multiple intermediate description texts obtained in the last round of error correction iteration when the error correction iteration stopping condition is met.

3. The method according to claim 2, characterized in that, The stopping condition for the error correction iteration is: The semantic difference rate of the multiple intermediate description texts obtained in the current round of error correction iteration is less than the semantic difference rate threshold; or, The number of error correction iterations reaches the threshold.

4. The method according to claim 2 or 3, characterized in that, The step involves performing at least one round of error correction iterations on the multiple initial description texts based on semantic differences, to obtain the multiple target description texts after error correction, including: In any round of error correction iteration, text semantic difference analysis is performed based on multiple descriptive texts to obtain the difference result corresponding to this round of error correction iteration. The difference result is used to indicate the semantic difference between every two descriptive texts in the multiple descriptive texts. Based on the difference results corresponding to the current error correction iteration and the multiple descriptive texts, error correction is performed to obtain multiple intermediate descriptive texts for the current error correction iteration; In the first round of error correction iteration, the plurality of description texts are the plurality of initial description texts; in the m-th round of error correction iteration, the plurality of description texts are the plurality of intermediate description texts obtained in the (m-1)-th round of error correction iteration.

5. The method according to claim 4, characterized in that, The step of performing text semantic difference analysis based on multiple descriptive texts to obtain the difference results corresponding to this round of error correction iteration includes: The multiple descriptive texts are input into the difference analysis model to obtain the difference results of every two descriptive texts in the multiple descriptive texts corresponding to the current error correction iteration.

6. The method according to claim 4, characterized in that, The multiple VLMs include a first VLM and a second VLM. The step of performing text semantic difference analysis based on multiple descriptive texts to obtain the difference results corresponding to this round of error correction iterations includes: The first description text and the second description text are input into the first VLM for text semantic difference analysis to obtain the first difference result corresponding to the current error correction iteration output by the first VLM. The first difference result is used to indicate the semantic difference between the first description text and the second description text. The difference result includes the first difference result. In the first round of error correction iteration, the first description text is the initial description text generated by the first VLM for the video, and the second description text is the initial description text generated by the second VLM for the video; in the m-th round of error correction iteration, the first description text and the second description text are the intermediate description texts obtained by error correction in the (m-1)-th round of error correction iteration. The step of performing error correction based on the difference results corresponding to the current error correction iteration and the multiple descriptive texts to obtain multiple intermediate descriptive texts for the current error correction iteration includes: Based on the first VLM, the first description text is corrected according to the first difference result to obtain the first intermediate description text corresponding to the current round of error correction iteration.

7. The method according to claim 6, characterized in that, The step of correcting the first description text based on the first VLM according to the first difference result to obtain the first intermediate description text corresponding to the current round of error correction iteration includes: A first error correction prompt word is generated based on the first difference result. The first error correction prompt word is used to guide the first VLM to correct the first description text. Input the video, the first description text, and the first error correction prompt into the first VLM; Based on the first VLM, the first description text is corrected to obtain the first intermediate description text after correction in this round of error correction iteration output by the first VLM.

8. The method according to claim 7, characterized in that, The first difference result is used to indicate different levels and types of semantic differences between the first descriptive text and the second descriptive text. The level includes at least one of the global, object, scene, target or action of the video, and the type includes at least one of the replacement difference, omission difference or addition difference. The replacement difference is used to indicate semantic differences that are described differently by the first description text compared to the second description text. The omission difference is used to indicate semantic differences that are not described by the first description text but are described by the second description text. The addition difference is used to indicate semantic differences that are described by the first description text but not by the second description text.

9. The method according to claim 8, characterized in that, The first difference result includes at least one difference item, each of the at least one difference item being used to indicate a semantic difference content of a type in a hierarchy; the generation of a first error correction prompt word based on the first difference result includes: Generate the first error correction prompt word corresponding to each of the difference items.

10. The method according to claim 9, characterized in that, The video, the first descriptive text, and the first error correction prompt are input into the first VLM; In the dialogue corresponding to each difference item in the first VLM, the video, the first descriptive text, each difference item, and the corresponding first error correction prompt word are input; The step of correcting the first description text based on the first VLM to obtain the first intermediate description text after correction in this round of error correction iteration, includes: Based on the first VLM, the first error correction prompt word corresponding to the video and the at least one difference item is used to correct the at least one difference item of the first description text, so as to obtain the first intermediate description text in the current error correction iteration output by the first VLM.

11. The method according to claim 10, characterized in that, In the dialogue corresponding to each difference item in the first VLM, the input of the video, the first descriptive text, each difference item, and the corresponding first error correction prompt word includes: Based on the time range corresponding to the target difference item in the local layer of the video, video segments corresponding to the time range are cropped from the video. The local layer refers to other layers besides the global layer. In the dialogue corresponding to the target difference item in the first VLM, the video clip, the first descriptive text, the target difference item, and the corresponding first error correction prompt word are input.

12. The method according to claim 11, characterized in that, The step of correcting at least one difference in the first description text based on the first VLM according to the first error correction prompt word corresponding to the video and the at least one difference item, to obtain the first intermediate description text in the current error correction iteration output by the first VLM, includes: Based on the first VLM, the first error correction prompt word corresponding to the target difference item and the video segment are used to correct the target difference item in the first descriptive text.

13. The method according to claim 12, characterized in that, The step of correcting the target difference item in the first descriptive text based on the first VLM according to the first correction prompt word corresponding to the target difference item and the video segment includes: Based on the target number of sampling frames per unit duration, the video segment is frame-sampling using the first VLM to obtain target video frames; the target number of sampling frames is greater than the number of sampling frames per unit duration that the first VLM performs frame-sampling on in the video. Based on the first VLM, the first error correction prompt word corresponding to the target difference item and the target video frame are used to correct the target difference item in the first description text.

14. The method according to claim 9, characterized in that, The step of inputting the video, the first descriptive text, and the first error correction prompt word into the first VLM includes: The video, the first descriptive text, the at least one difference item and the corresponding first error correction prompt word, and the time range corresponding to the target difference item at the local level in the video are input into a round of dialogue of the first VLM, wherein the local level is other levels besides the global level.

15. The method according to any one of claims 1-14, characterized in that, The step of performing text semantic fusion on the multiple target description texts to obtain the fused description text includes: The multiple target description texts are input into a large language model (LLM), and semantic fusion is performed based on the large language model to obtain the fused description text.

16. The method according to any one of claims 1-14, characterized in that, The step of performing text semantic fusion on the multiple target description texts to obtain the fused description text includes: The multiple target description texts are semantically fused using a weighted fusion algorithm to obtain the fused description text.

17. A description of a text correction device, characterized in that, The device includes: The acquisition module is used to acquire multiple initial description texts generated by multiple VLMs for the video; The error correction module is used to correct the multiple initial description texts based on the semantic differences between them, so as to obtain multiple target description texts after error correction. The fusion module is used to perform text semantic fusion on the multiple target description texts to obtain the fused description text.

18. An electronic device, characterized in that, The electronic device includes a processor and a memory; the processor is coupled to the memory; the memory is used to store computer instructions, which are loaded and executed by the processor to enable the electronic device to perform the method as described in any one of claims 1-16.

19. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-16.

20. A computer program product containing instructions, characterized in that, When the instruction is executed by an electronic device or a cluster of computing devices, it causes the electronic device or the cluster of computing devices to perform the method as described in any one of claims 1-16.

21. A computer-readable storage medium, characterized in that, It includes computer program instructions that, when executed by an electronic device or a cluster of computing devices, perform the method as described in any one of claims 1-16.

22. A chip system, characterized in that, The chip system includes a memory and a processor, the memory being used to store a set of computer instructions, and when the processor executes the set of computer instructions, it performs the method as described in any one of claims 1-16.