Data generation method, model optimization method, and related device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2024-12-02
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]鉴于以上内容,有必要提供一种数据生成方法、模型优化方法及相关设备,能够解决模型生成长视频的文本描述时容易出现幻觉导致的难以得到长视频的准确文本描述的问题
[0017]通过上述技术方案,可以基于幻觉判断提示生成多种幻觉判断问题,基于幻觉判断问题从多个维度确定初始视频级描述中的幻觉,从而为后续的幻觉优化提供基础。基于幻觉优化提示优化初始视频级描述中的幻觉,自然语言模型最终输出的优化的视频级描述相较于初始视频级描述中的幻觉数量很少,从而可以使用优化的视频级描述与对应的优化的视频级描述作为数据对,以实现对自然语言模型的训练。
Smart Images

Figure CN120431502B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence and relates to a data generation method, a model optimization method, and related equipment. Background Technology
[0002] With the rapid development of multimodal large language models, they can now automatically generate corresponding text descriptions for videos. However, due to the large amount of information and complex scenes contained in long videos, multimodal large language models are prone to producing inaccurate or unrealistic descriptions when processing them, leading to an "illusion" phenomenon in the descriptive text. While methods that segment and reassemble descriptions of long videos can alleviate the illusion problem, repeated use of the model can still lead to the accumulation of illusions, making it difficult to obtain accurate text descriptions of long videos. Summary of the Invention
[0003] In view of the above, it is necessary to provide a data generation method, a model optimization method, and related equipment that can solve the problem of difficulty in obtaining accurate text descriptions of long videos due to the illusion that easily occurs when the model generates text descriptions of long videos.
[0004] Firstly, this application provides a data generation method based on hallucination enhancement, applied to electronic devices. The method includes: segmenting a video to obtain multiple video segments; obtaining a segment-level description corresponding to each video segment; and updating the multiple segment-level descriptions based on preset first prompt information to obtain multiple segment-level descriptions including hallucinations. Through the above technical solution, hallucinations generated during the video description generation process can be taken into account. A dataset is constructed to address the hallucinations that occur, allowing for the training of a model to reduce or avoid hallucinations during video description generation, thereby providing a data foundation for training the model.
[0005] In one possible implementation, segmenting the video to obtain multiple video segments includes: segmenting the video into scenes based on the scene similarity between adjacent video frames in the video to obtain segmented videos corresponding to each scene; and splicing the segmented videos based on preset duration parameters to obtain the video segments.
[0006] The above technical solution allows a video to be segmented into multiple independent segments, each focusing on a specific scene or plot, thus facilitating subsequent analysis and description. Merging multiple consecutive segmented videos into a single, longer video segment not only reduces the number of subsequent video descriptions, improving processing efficiency, but also avoids the accumulation of excessive hallucinations in the video description, thereby increasing the accuracy of the generated video description.
[0007] In one possible implementation, updating the data of multiple paragraph-level descriptions based on preset first prompt information includes: updating at least one data for each paragraph-level description based on multiple hallucination generation prompts in the first prompt information, such that each paragraph-level description updated contains at least one hallucination corresponding to a hallucination generation prompt.
[0008] In one possible implementation, the multiple illusion generation prompts include a first illusion generation prompt and a second illusion generation prompt, wherein the first illusion generation prompt includes prompt information for causing the language expression of each paragraph-level description to generate an illusion, and the second illusion generation prompt includes prompt information for causing the entity object of each paragraph-level description to generate an illusion.
[0009] In one possible implementation, the first illusion generation prompt includes one or more of a logical illusion generation prompt and a redundant illusion generation prompt, wherein: the logical illusion generation prompt is used to instruct the random shuffling of the text order of each paragraph-level description; and the redundant illusion generation prompt is used to instruct the random repetition of the text statements of each paragraph-level description.
[0010] In one possible implementation, the second illusion generation prompt includes one or more of entity illusion generation prompts, attribute illusion generation prompts, quantity illusion generation prompts, and behavior illusion generation prompts, wherein: the entity illusion generation prompt is used to instruct the entity object of each paragraph-level description to be updated; the attribute illusion generation prompt is used to instruct the adjective of each paragraph-level description to be updated; the quantity illusion generation prompt is used to instruct the quantifier of each paragraph-level description to be updated; and the behavior illusion generation prompt is used to instruct the verb of each paragraph-level description to be updated.
[0011] The above technical solution can generate illusions in multiple dimensions for paragraph-level descriptions, so that the updated paragraph-level descriptions contain illusion data in multiple dimensions. This allows the model to be trained in subsequent processes to learn to avoid generating illusion data and improve the accuracy of paragraph-level description generation.
[0012] Secondly, this application provides a data generation method based on hallucination optimization, applied to electronic devices. The method includes: aggregating multiple paragraph-level descriptions containing hallucinations to obtain an initial video-level description; and optimizing the initial video-level description based on preset second prompt information to obtain an optimized video-level description.
[0013] The above technical solution can take into account the hallucinations that may occur during the aggregation video description process. A dataset can be constructed to train the model to reduce or avoid hallucinations during the aggregation video description process, thereby providing a data foundation for training the model.
[0014] In one possible implementation, optimizing the initial video-level description based on preset second prompt information to obtain an optimized video-level description includes: determining the presence of a hallucination in the initial video-level description based on the hallucination judgment prompt in the second prompt information and the multiple segment-level descriptions; and optimizing the hallucination based on the hallucination optimization prompt in the second prompt information to obtain the optimized video-level description.
[0015] In one possible implementation, the illusion optimization prompt includes an optimization reason prompt, and the method further includes: generating a reason for optimizing the illusion based on the optimization reason prompt.
[0016] In one possible implementation, the illusion optimization prompt further includes an optimization content prompt, and the method further includes: obtaining the optimization content for the illusion based on the optimization content prompt.
[0017] The above technical solution allows for the generation of various hallucination judgment questions based on hallucination judgment prompts. These questions then help determine the hallucinations in the initial video-level description from multiple dimensions, providing a foundation for subsequent hallucination optimization. The optimized video-level descriptions output by the natural language model are then optimized based on hallucination optimization prompts. The number of hallucinations in the final optimized video-level description is significantly less than that in the initial description, allowing the optimized video-level descriptions and their corresponding optimized video-level descriptions to be used as data pairs for training the natural language model.
[0018] Thirdly, this application provides a model optimization method applied to electronic devices. The method includes: obtaining corresponding segment-level descriptions based on video segments using a preset video description model; obtaining segment-level descriptions including hallucinations corresponding to the segment-level descriptions, wherein the segment-level descriptions including hallucinations are obtained according to the aforementioned hallucination-enhanced data generation method; and performing supervised optimization on the video description model based on the segment-level descriptions and the segment-level descriptions including hallucinations.
[0019] The above technical solution employs a supervised learning method, inputting the original paragraph-level descriptions and corresponding paragraph-level descriptions, including those involving hallucinations, as paired training data into the video description model. This allows the model to learn and distinguish between the two. Through iterative training and optimization, the video description model gradually improves its accuracy in identifying hallucination descriptions until it reaches a preset threshold. This enables the video description model to generate accurate paragraph-level descriptions and effectively avoid hallucinations.
[0020] Fourthly, this application provides a model optimization method applied to electronic devices. The method includes: aggregating paragraph-level descriptions using a preset natural language processing model to obtain an initial video-level description; obtaining an optimized video-level description corresponding to the initial video-level description, wherein the optimized video-level description is obtained according to the aforementioned data generation method based on illusion optimization; and performing supervised optimization on the natural language processing model based on the initial video-level description and the optimized video-level description.
[0021] The above technical solution employs a supervised learning method, inputting initial video-level descriptions and corresponding optimized video-level descriptions as paired training data into the natural language processing model. This enables the model to learn and distinguish between the two. Through iterative training and optimization, the natural language processing model gradually improves its accuracy in judging hallucination descriptions until it reaches a preset threshold. This allows the natural language processing model to aggregate paragraph-level descriptions into accurate video-level descriptions and effectively avoid hallucinations.
[0022] Fifthly, this application provides an electronic device, which includes a memory and a processor: wherein the memory is used to store program instructions; and the processor is used to read and execute the program instructions stored in the memory, wherein when the program instructions are executed by the processor, the electronic device performs the aforementioned data generation method and model optimization method.
[0023] Sixthly, this application provides a chip coupled to a memory in an electronic device, the chip being used to control the processor of the electronic device to execute the aforementioned data generation method and model optimization method.
[0024] In a seventh aspect, this application provides a computer storage medium storing program instructions that, when executed on an electronic device, cause the processor of the electronic device to perform the aforementioned data generation method and model optimization method.
[0025] Furthermore, the technical effects brought about by aspects five through seven can be found in the descriptions of the methods in the above-mentioned method section, and will not be repeated here. Attached Figure Description
[0026] Figure 1 This is an example diagram illustrating how hallucinations can cause application problems in a video description provided in an embodiment of this application.
[0027] Figure 2 This is an example diagram illustrating how hallucinations can cause application problems in a video description provided in another embodiment of this application.
[0028] Figure 3 This is an example diagram of the data generation process of the related technology provided in an embodiment of this application.
[0029] Figure 4 This is an example diagram of the data generation process of the related technology provided in another embodiment of this application.
[0030] Figure 5 This is a software architecture diagram of an electronic device provided in an embodiment of this application.
[0031] Figure 6 This is a flowchart of a data generation method based on illusion enhancement provided in an embodiment of this application.
[0032] Figure 7 This is a flowchart of a video segmentation method provided in an embodiment of this application.
[0033] Figure 8 This is an example diagram of segmenting a video and video segments provided in an embodiment of this application.
[0034] Figure 9 This is a flowchart of a model optimization method provided in an embodiment of this application.
[0035] Figure 10 This is a flowchart of a data generation method based on illusion optimization provided in an embodiment of this application.
[0036] Figure 11 This is a flowchart of a model optimization method provided in another embodiment of this application.
[0037] Figure 12 This is a flowchart of a data generation method based on illusion enhancement and optimization provided in an embodiment of this application.
[0038] Figure 13 This is a schematic diagram of the functional modules of a data generation method provided in an embodiment of this application.
[0039] Figure 14 This is a hardware architecture diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0040] In one embodiment of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in one embodiment of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used in this application's specification is for the purpose of describing particular embodiments only and is not intended to limit the application. It should be understood that, unless otherwise stated, " / " in this application means "or". For example, A / B can mean A or B. "And / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. "At least one" refers to one or more. "More than one" refers to two or more. For example, at least one of a, b, or c can represent: a, b, c, a and b, a and c, b and c, and a, b, and c. Where there is no conflict, the following embodiments and features described herein can be combined with each other.
[0042] With the rapid development of multimodal large language models, they have demonstrated powerful capabilities, automatically generating corresponding text descriptions for video content. This function has greatly promoted the development of the field of video understanding, enabling machines to interpret information in videos more intuitively. However, multimodal large language models face a series of challenges when processing long videos.
[0043] Long videos typically contain a large amount of information and complex scenes, placing high demands on the processing power of the model. Due to the limitations of the model, the descriptive text for long videos often exhibits inaccurate descriptions or "illusions" that deviate from the actual content of the video. These illusory descriptions not only mislead users' understanding of the video content but may also lead to a series of problems in practical applications.
[0044] In one example, given search text containing keywords related to a specific scene, accurate video descriptions can help users quickly locate videos containing that scene from multiple videos. However, if the video description contains illusions, the accuracy of scene retrieval will be affected, potentially making it difficult for users to quickly and accurately locate the video corresponding to their search query.
[0045] For example Figure 1 The image shown is an example of an application problem caused by hallucinations in a video description provided in an embodiment of this application. In this case, the user... Figure 1 The interactive interface shown in Figure (a) inputs search text for a specific scenario, which includes "Leo and the cat". Based on the search text, [the system can]... Figure 1 The multiple videos shown in Figure (b) are filtered to select those whose video descriptions match the search text, thus allowing users to quickly find and present the videos they need. For example, the videos found and presented include... Figure 1As shown in Figure (c), video A has a video description of "Leo walking a kitten," containing the keywords "Leo" and "cat" that correspond to the search text. However, the keyword "cat" is an illusion in the video description of video A; the actual scene in video A is "Leo walking a puppy." Therefore, the specific scene the user is looking for does not exist in video A, resulting in a search error caused by an illusion.
[0046] In another example, generating corresponding text descriptions for videos can provide real-time navigation for visually impaired individuals. For instance, by recording a video and obtaining a real-time scene description, visually impaired individuals can navigate more easily based on the speech information derived from that description. However, if the video description contains illusions, it could lead to errors in the description of key scenes, thereby misleading visually impaired individuals and posing safety hazards to their travel.
[0047] For example Figure 2 The image shown is an example of an application problem caused by a hallucination in a video description provided in another embodiment of this application. In this illustration, the intelligent assistant "Xiao Zhu" responds to the user's request, "Xiao Zhu, I'm out, please keep an eye on the road," by activating a safety mode, acquiring real-time video of the user's environment, and outputting a video description: "Beware of the three steps not far ahead, please be careful." However, the "three steps" described are a hallucination in the video description; the actual number of steps in the environment is four. This hallucination in the video description may pose a safety hazard for visually impaired individuals.
[0048] To mitigate the hallucination problem in video descriptions of long videos generated by models, related techniques employ a segmented description and then aggregation approach. For example, a long video is cut into multiple shorter segments, each segment is described independently, and these descriptions are then aggregated to form a complete video description. While this method can alleviate the hallucination problem to some extent, repeatedly using the model for segmented descriptions still struggles to prevent hallucinations from occurring and may even lead to their accumulation, resulting in a final aggregated description with even more hallucinations.
[0049] In summary, current multimodal large language models suffer from a significant illusion problem in long video descriptions, which severely impacts the accuracy and practicality of video descriptions.
[0050] The relevant technologies do not provide a data generation method that can effectively solve the problem of hallucinations. (Reference) Figure 3 The diagram shown is an example of a data generation process for a related technology provided in an embodiment of this application. Figure 3 The image shown is from The AI-driven video recap data generation process first generates segment-level descriptions for the video, then uses a Large Language Model (LLM) to aggregate them into video-level descriptions for training data generation. Specifically, a 60-minute complete video is broken down into multiple 5-minute video segments, and a segment-level description is determined for each segment. For example, the segment-level description for the first video segment might be "C opens the refrigerator," the second segment "C peels an onion," the third segment "C cuts an onion," and the last segment "C turns off the stove." The LLM then aggregates these segment-level descriptions to obtain the video-level description: "C takes an onion out of the refrigerator and cuts it." The video-level description is a pseudo-segment description.
[0051] refer to Figure 4 The diagram shown is an example of a data generation process for a related technology provided in another embodiment of this application. Figure 4 The image shown is from Stanford. and Google The VIDEO-STaR data generation process utilizes a large language model to determine the quality and confidence of generated descriptions. Starting with input videos containing video frame labels, the model is optimized through instruction tuning. Based on the model, questions and answers corresponding to the videos are generated. The accuracy of the model's answers is determined by comparing them with the video frame labels, and the model is then optimized in reverse based on the answer accuracy. For example... Figure 4 As shown, the question corresponding to the video generated based on the model is: "What is the action sequence in the video?". The corresponding answer is: "First, we see the man bending down and lifting a book. Then,... Finally, the man can be seen reading a book." Since the answer contains the video frame labels "Bend down" and "Read Book," it can be determined that the answer is consistent with the video frame labels, and the model's answer to the above question is accurate.
[0052] Although the methods mentioned above can generate training data for video description using multimodal large language models, they do not take into account the hallucinations present when generating video description data using large models, nor do they construct corresponding training data specifically for the hallucination problem. Therefore, they cannot avoid the impact of hallucinations on the accuracy of descriptions, nor can they train models to optimize the hallucination problem.
[0053] To address the aforementioned issues, this application provides a data generation method for generating data to address the problem of optimizing hallucinations. This method can construct training data for hallucination problems, thereby enabling the optimization of models based on the training data to avoid hallucination problems.
[0054] The data generation method provided in this application is applied to electronic devices, such as mobile phones, tablets, wearable devices, camera devices, computers, and self-moving devices. The following will combine... Figure 5 This application provides a software architecture diagram of the electronic device provided in its embodiments.
[0055] See Figure 5 As shown, the layered architecture in electronic devices divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. For example, the Android system, from top to bottom, consists of the application layer 101, framework layer 102, Android runtime and system libraries 103, hardware abstraction layer 104, kernel layer 105, and hardware layer 106.
[0056] Application layer 101 may include a series of application packages. For example, application packages may include applications such as camera, gallery, calendar, calling, map, navigation, WLAN, Bluetooth, music, video, SMS, device control services, etc.
[0057] The framework layer 102 provides an Application Programming Interface (API) and programming framework for applications in the application layer. The application framework layer includes predefined functions. For example, it may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0058] The window manager manages window programs. It can obtain screen size, determine the presence of a status bar, lock the screen, and capture screenshots. The content provider stores and retrieves data, making it accessible to applications. This data can include videos, images, audio, made and received calls, browsing history and bookmarks, phone books, etc. The view system includes visual controls, such as controls for displaying text and controls for displaying images. The view system can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon can include views for displaying text and views for displaying images. The phone manager provides communication functionality for electronic devices, such as managing call status (including connection and disconnection). The resource manager provides applications with various resources, such as localized strings, icons, images, layout files, and video files. The notification manager allows applications to display notifications in the status bar, conveying informational messages that disappear automatically after a short pause without user interaction. For example, the notification manager is used to notify of download completion or message alerts. The notification manager can also display notifications as icons or scrolling text in the system's top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting alert sounds, causing electronic devices to vibrate, and flashing indicator lights.
[0059] The Android Runtime consists of the core libraries and the virtual machine. The Android runtime is responsible for the scheduling and management of the Android system. The core libraries consist of two parts: one part contains the functionalities that the Java language needs to call, and the other part contains the core Android libraries.
[0060] Application layer 101 and framework layer 102 run in a virtual machine. The virtual machine executes the Java files of the application layer and framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0061] System library 103 may include multiple functional modules. For example, a surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0062] The Surface Manager manages the display subsystem and provides fusion of 2D and 3D layers for multiple applications. The Media Library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG. The 3D Graphics Processing Library implements 3D graphics drawing, image rendering, compositing, and layer processing. The 2D Graphics Engine is the drawing engine for 2D graphics.
[0063] Hardware Abstraction Layer 104 runs in user space, encapsulates kernel-level drivers, and provides calling interfaces to the upper layers.
[0064] Kernel layer 105 is the layer between hardware and software. Kernel layer 105 contains at least the display driver, touch driver, audio driver, and sensor driver.
[0065] The kernel layer (105) is the core of the operating system for electronic devices. It is the first layer of software extension based on the hardware, providing the most basic functions of the operating system. It is the foundation for the operation of the operating system, responsible for managing system processes, memory, device drivers, files, and network systems, and determining the system's performance and stability. For example, the kernel layer can determine the timing of an application's operation on a certain part of the hardware.
[0066] Kernel layer 105 includes hardware-dependent programs such as interrupt handlers and device drivers, as well as basic, common, and frequently running modules such as clock management and process scheduling modules, and critical data structures. The kernel layer can be located within the processor or embedded in internal memory.
[0067] Hardware layer 106 includes the hardware of electronic devices, such as displays, buttons, cameras, etc.
[0068] See Figure 6 The diagram shown is a flowchart of a data generation method based on illusion enhancement provided in an embodiment of this application. The method is applied in an electronic device and includes the following steps.
[0069] S601 divides the video into multiple video segments.
[0070] In one embodiment of this application, if the video is long, for example, longer than a preset duration (such as 2 minutes, 10 minutes, etc.), analyzing or describing the video as a whole is not only time-consuming and laborious, but may also make it difficult to accurately capture the key information and details in the video. By segmenting the video, a lengthy video can be broken down into multiple video segments (also called "video clips") that are easier to process, which helps to improve the efficiency and accuracy of video description.
[0071] In one embodiment of this application, when segmenting a video, the video can be divided into multiple segmented videos containing a single scene based on the similarity of scenes between video frames. Since the duration of the segmented videos may be short (e.g., less than 30 seconds), the number of segmented videos may be relatively large. Multiple consecutive segmented videos can be merged into a single longer video segment, which not only reduces the number of subsequent video descriptions, thus improving processing efficiency, but also avoids the excessive accumulation of hallucinations in the video description, improving the accuracy of the generated video description. In one example, the method for segmenting the video can be referred to below. Figure 7 Description of the illustrated embodiment.
[0072] The data generation method provided in this application embodiment can be used to construct a dataset for training a model. Therefore, the video segmented in step S601 can include multiple videos. The data generation method provided in this application embodiment is executed on each video to generate a dataset containing a rich amount of data. The number of videos can be determined according to actual needs, and this application does not impose specific limitations on this.
[0073] S602, obtain the paragraph-level description corresponding to each video segment.
[0074] In one embodiment of this application, a preset video description model can be used to obtain the paragraph-level description corresponding to each video segment. In one example, the video description model can be a multimodal large language model, which has the ability to process video and text data and can generate descriptive text for video segments as paragraph-level descriptions. For example, the multimodal large language model can include, but is not limited to, InternVL2.
[0075] In one embodiment of this application, a video segment and a preset first prompt word can be input into a multimodal large language model. The multimodal large language model analyzes the video segment based on the first prompt word and outputs a segment-level description of the video segment. The first prompt word can be information (e.g., a phrase or sentence) used to guide the model in generating a segment-level description of the video segment.
[0076] In one example, the paragraph-level description 1 of video segment 1 could be exemplarily represented as: "This video clip shows a herd of elephants traversing a vast, dusty grassland. The elephants, including a mother and her calf, move steadily through the desolate landscape, each step kicking up a cloud of dust. The scene is set under a hazy sky, highlighting the arid environment. Close-up shots capture the elephants' large ears and the details of their skin, emphasizing their majestic presence. The calf follows closely behind its mother, symbolizing the bond and protection within the group. The overall atmosphere is one of these magnificent creatures' indomitable journey through their natural habitat."
[0077] The paragraph-level description 2 for video segment 2 can be exemplarily represented as: "In this clip of the video, a baby elephant interacts with its mother in a tranquil natural environment. The scene captures a heartwarming moment as the mother elephant gently touches the baby elephant's head with her trunk, conveying love and affection. The background is a vast grassland, with sparse vegetation highlighting the immensity of the African landscape. Soft light falls on the scene, emphasizing the fine texture of the elephant's skin and the tender bond between them. The overall atmosphere is serene, showcasing the deep connection between the mother and baby elephant."
[0078] The paragraph-level description of video segment 3 could be exemplarily stated as: "This segment presents a stunning aerial view overlooking a vast landscape. The scene is dominated by the stark contrast between the dark waters of a large lake and the surrounding arid, reddish-brown terrain. Numerous small, circular formations, possibly from saline soil or mineral deposits, dot the lake's surface, creating mesmerizing patterns. Soft pastel hues in the sky add to the surreal beauty of the landscape. A lone elephant can be seen walking in the distance, emphasizing the scale and isolation of the environment. The overall atmosphere is one of desolate yet transcendent beauty, with natural elements harmoniously blending to create a captivating visual experience."
[0079] In the above example, the generated paragraph-level descriptions may contain illusory data, but this does not affect the data generation in subsequent processes. For example, in paragraph-level description 1, "Elephants, including a mother elephant and her calf, move steadily across a desolate landscape, raising a cloud of dust with each step" is somewhat redundant with the previous statement, "The video clip shows a herd of elephants crossing a vast, dusty grassland."
[0080] Through the above embodiments, a segment-level description corresponding to each video segment can be obtained based on a multimodal large language model, thereby providing a data foundation for the subsequent hallucination generation process.
[0081] S603, based on the preset first prompt information, update the data of multiple paragraph-level descriptions to obtain multiple paragraph-level descriptions including hallucinations.
[0082] In one embodiment of this application, the preset first prompt information includes multiple hallucination generation prompts. Based on the multiple hallucination generation prompts, at least one data update is performed on each paragraph-level description, such that each updated paragraph-level description contains at least one hallucination corresponding to a hallucination generation prompt. In this way, hallucination generation can be performed on the paragraph-level description from multiple dimensions, resulting in an updated paragraph-level description containing more hallucinations.
[0083] In one embodiment of this application, paragraph-level descriptions and various hallucination generation prompts can be input into a pre-trained natural language processing (NLP) model. The NLP model then generates corresponding paragraph-level descriptions including hallucinations based on these prompts. For example, the NLP model may include, but is not limited to, Llama3 (Large Language Model Architecture 3.0). The following will illustrate, with several examples, how to obtain corresponding paragraph-level descriptions including hallucinations based on various hallucination generation prompts.
[0084] In one embodiment of this application, based on prior experience, the criteria for judging the accuracy of paragraph-level descriptions typically include: the accuracy of the language expression and the accuracy of the description of entity objects (such as people, animals, buildings, plants, scenes, etc.) in the video. Therefore, based on the characteristics of the descriptive content and processing requirements, illusion generation prompts can be designed for both language expression and entity object descriptions. The following embodiments illustrate this from different perspectives using various illusion generation prompts (e.g., first illusion generation prompts and second illusion generation prompts).
[0085] In one embodiment of this application, the first illusion generation prompt includes prompt information for generating illusions in the language expressions of each paragraph-level description. The first illusion generation prompt includes one or more of a logical illusion generation prompt and a redundant illusion generation prompt, wherein: the logical illusion generation prompt is used to instruct the random shuffling of the text order of each paragraph-level description; the redundant illusion generation prompt is used to instruct the random repetition of the text statements of each paragraph-level description.
[0086] In one example, the logical illusion generation prompt can be exemplarily represented as: "Swap the positions of two sentences in the paragraph-level description to increase the illusion of logical order, while keeping other content unchanged." Based on the logical illusion generation prompt, the natural language processing model can randomly adjust the text order in the paragraph-level description, thereby obtaining a paragraph-level description that includes logical illusion. The number of text sentences swapped can be set according to actual needs; the above is merely an example, and this application does not limit this.
[0087] For example, based on the logic illusion generation prompt, the last two sentences of paragraph-level description 1 were swapped, resulting in paragraph-level description 11, which includes the logic illusion: "This video clip shows a herd of elephants traversing a vast, dusty grassland. The elephants, including a mother and her calf, move steadily through the desolate landscape, each step kicking up a cloud of dust. The scene is set under a gloomy sky, highlighting the arid environment. Close-up shots capture the elephants' large ears and the fine texture of their skin, emphasizing their majestic presence. The overall atmosphere is one of these magnificent creatures' indomitable journey in their natural habitat. The calf follows closely behind its mother, symbolizing the bond and protection within the group."
[0088] By setting logical illusion generation prompts, paragraph-level descriptions containing logical illusions can be generated, achieving the generation of illusions in the language expression of paragraph-level descriptions and increasing the diversity of illusionary content in language expression. Based on paragraph-level descriptions containing logical illusions, the model can learn to avoid generating logical illusions in the generated paragraph-level descriptions in subsequent processes.
[0089] In one example, the redundancy illusion generation prompt can be exemplified as: "Repeat a single sentence in the paragraph-level description to add redundant illusionary information, while keeping other content unchanged." Based on the logical illusion generation prompt, the natural language processing model can randomly repeat text statements in the paragraph-level description, thereby obtaining a paragraph-level description that includes logical illusions. The number of repeated text statements can be set according to actual needs, and this application does not limit this.
[0090] For example, based on the redundancy illusion generation prompt, the sentence "The scene is set under a gloomy sky, highlighting the arid environment" in paragraph-level description 1 was repeated, resulting in paragraph-level description 12, which includes the redundancy illusion: "This video clip shows a herd of elephants traversing a vast, dusty grassland. The elephants, including a mother and her calf, move steadily through the desolate landscape, each step kicking up a cloud of dust. The scene is set under a gloomy sky, highlighting the arid environment. The elephants' large ears and the fine texture of their skin are captured in close-up shots, emphasizing their majestic presence. The calf follows closely behind its mother, symbolizing the bond and protection within the group. The overall atmosphere is one of these magnificent creatures' indomitable journey in their natural habitat."
[0091] By setting redundancy illusion generation prompts, paragraph-level descriptions containing redundancy illusions can be generated, achieving illusion generation of language expressions at the paragraph level and increasing the diversity of illusion content in language expressions. Based on paragraph-level descriptions containing redundancy illusions, the model can learn to avoid generating redundant illusions in the generated paragraph-level descriptions in subsequent processes.
[0092] In one embodiment of this application, the second illusion generation includes prompting information for generating illusions for each paragraph-level described entity object. The second illusion generation prompts include one or more of entity illusion generation prompts, attribute illusion generation prompts, quantity illusion generation prompts, and behavior illusion generation prompts.
[0093] In one embodiment of this application, the entity illusion generation prompt is used to instruct the entity objects in each paragraph-level description to be updated, such as updating the category of the entity objects. In one example, the entity illusion generation prompt can be exemplarily represented as: "Modify only a single sentence in the paragraph-level description by replacing the object in the sentence with a similar but different category object, keeping other content unchanged." For example, a mapping table between similar but different categories of objects can be pre-constructed. The natural language processing model can identify objects in the paragraph-level description, search for objects in the mapping table that are similar to the identified objects but of a different category, and replace the identified objects with the found objects to achieve entity illusion generation.
[0094] For example, by replacing the object "elephant" in paragraph-level description 1 with "giraffe" based on the entity illusion generation prompt, the resulting paragraph-level description 13, which includes entity illusion, is: "This video clip shows a herd of giraffes traversing a vast, dusty grassland. The giraffes, including a mother giraffe and her calf, move steadily through the desolate landscape, each step kicking up a cloud of dust. The scene is set under a gloomy sky, highlighting the arid environment. Close-up shots capture the giraffes' large ears and the fine texture of their skin, emphasizing their majestic presence. The calf follows closely behind its mother, symbolizing the bond and protection within the group. The overall atmosphere is one of these magnificent creatures' indomitable journey in their natural habitat."
[0095] By setting entity illusion generation prompts, paragraph-level descriptions containing entity illusions can be generated, achieving the generation of entity illusions for paragraph-level descriptions and increasing the diversity of entity illusion content. Based on paragraph-level descriptions containing entity illusions, the model can learn how to avoid entity illusions in the generated paragraph-level descriptions in subsequent processes.
[0096] In one embodiment of this application, the attribute illusion generation prompt is used to instruct the adjectives of each paragraph-level description to be updated. In one example, the attribute illusion generation prompt can be exemplarily represented as: "Modify only a single sentence in the paragraph-level description by replacing the adjectives of the objects in that sentence, keeping other content unchanged." For example, an adjective list can be pre-constructed, which may contain adjectives with different or opposite meanings. The natural language processing model can identify the adjectives of the objects in the paragraph-level description, select adjectives with different or opposite meanings from the adjective list that have different or opposite meanings from the identified adjectives, and replace the identified adjectives based on the selected adjectives, thereby generating the attribute illusion.
[0097] For example, by replacing "big ears" with "small ears" in paragraph-level description 1 based on attribute illusion generation prompts, the resulting paragraph-level description 14, which includes attribute illusion, is: "This video clip shows a herd of elephants traversing a vast, dusty grassland. The elephants, including a mother and her calf, move steadily through the desolate landscape, each step kicking up a cloud of dust. The scene is set under a gloomy sky, highlighting the arid environment. Close-up shots capture the small ears and the fine texture of the elephants' skin, emphasizing their majestic presence. The calf follows closely behind its mother, symbolizing the bond and protection within the group. The overall atmosphere is one of these magnificent creatures' indomitable journey in their natural habitat."
[0098] By setting attribute illusion generation prompts, paragraph-level descriptions containing attribute illusions can be generated, achieving illusion generation of entity objects for paragraph-level descriptions and increasing the diversity of illusion content for entity objects. Based on paragraph-level descriptions containing attribute illusions, the model can learn how to avoid attribute illusions in the generated paragraph-level descriptions in subsequent processes.
[0099] In one embodiment of this application, a quantity illusion generation prompt is used to instruct the updating of quantifiers for each paragraph-level description. In one example, the quantity illusion generation prompt can be exemplarily represented as: "Modify only a single sentence in the paragraph-level description, changing the quantity of the object while keeping other content unchanged." For example, a quantifier list can be pre-constructed, which may contain quantifiers with different meanings. The natural language processing model can identify the quantifiers of objects in the paragraph-level description, select quantifiers from the quantifier list that have different meanings from the identified quantifiers, and replace the identified quantifiers based on the selected quantifiers, thereby generating a quantity illusion.
[0100] For example, by replacing "a herd of elephants" with "a single elephant" in paragraph-level description 1 based on the quantity illusion generation prompt, the resulting paragraph-level description 15, which includes the quantity illusion, is: "This video clip shows an elephant traversing a vast, dusty grassland. The elephant, including a mother and her calf, moves steadily through the desolate landscape, raising a cloud of dust with each step. The scene is set under a gloomy sky, highlighting the arid environment. Close-up shots capture the elephant's large ears and the fine texture of its skin, emphasizing its majestic presence. The calf follows closely behind its mother, symbolizing the bond and protection within the group. The overall atmosphere is one of this magnificent creature's indomitable journey through its natural habitat."
[0101] By setting quantity illusion generation prompts, paragraph-level descriptions containing quantity illusions can be generated, achieving the illusion generation of entity objects in paragraph-level descriptions and increasing the diversity of the illusion content of entity objects. Based on paragraph-level descriptions containing quantity illusions, the model can learn to avoid generating quantity illusions in the generated paragraph-level descriptions in subsequent processes.
[0102] In one embodiment of this application, the behavior illusion generation prompt is used to instruct the verbs of each paragraph-level description to be updated. In one example, the behavior illusion generation prompt can be exemplarily represented as: "Modify only a single sentence in the paragraph-level description by replacing the action with its opposite action, keeping other content unchanged." For example, a verb list can be pre-constructed, which may contain verbs with different or opposite meanings. The natural language processing model can identify the verbs of objects in the paragraph-level description, select verbs from the verb list that have different or opposite meanings from the identified verbs, and replace the identified verbs of objects based on the selected verbs, thereby generating the behavior illusion.
[0103] For example, by replacing "moving steadily" with "remaining still" in paragraph-level description 1 based on behavioral hallucination generation prompts, paragraph-level description 16, which includes behavioral hallucination, is: "This video clip shows a herd of elephants crossing a vast, dusty grassland. The elephants, including a mother and her calf, stand motionless in the desolate landscape, without raising any dust. The scene is set under a gloomy sky, highlighting the arid environment. Close-up shots capture the elephants' large ears and the fine texture of their skin, emphasizing their majestic presence. The calf follows closely behind its mother, symbolizing the bond and protection within the group. The overall atmosphere is one of these magnificent creatures' indomitable journey in their natural habitat."
[0104] By setting behavioral illusion generation prompts, paragraph-level descriptions containing behavioral illusions can be generated, achieving illusion generation of entity objects within paragraph-level descriptions and increasing the diversity of illusionary content for entity objects. Based on paragraph-level descriptions containing behavioral illusions, the model can learn to avoid generating behavioral illusions in the generated paragraph-level descriptions in subsequent processes.
[0105] Through the above embodiments, illusion generation can be performed on paragraph-level descriptions in multiple dimensions, so that the updated paragraph-level descriptions contain illusion data in multiple dimensions. This allows the model to be trained in subsequent processes to learn to avoid generating illusion data and improve the accuracy of paragraph-level description generation.
[0106] In one embodiment of this application, the original paragraph-level description (e.g., the paragraph-level description obtained in S602) and the corresponding paragraph-level description including hallucinations can be used as a pair of training data. The video description model is trained using the paired training data with a supervised learning method, so that the model learns to distinguish between the original paragraph-level description and the corresponding paragraph-level description including hallucinations. This enables the trained video description model to generate accurate paragraph-level descriptions, thereby avoiding hallucinations in description.
[0107] In one example, the original paragraph-level description contains illusions, but the number of illusions in the original paragraph-level description is much smaller compared to the actively generated paragraph-level description that includes illusions. Therefore, using the original paragraph-level description and the corresponding paragraph-level description that includes illusions as data pairs, the video description model can be trained.
[0108] In another example, the original paragraph-level descriptions can be optimized to reduce illusions within them, and the optimized paragraph-level descriptions can be used for data generation in subsequent processes. Thus, the illusion-optimized paragraph-level descriptions and their corresponding illusion-included paragraph-level descriptions can be used as data pairs, further improving the model's training efficiency and effectiveness.
[0109] The hallucination-enhanced data generation method provided in this application generates diverse hallucination segment descriptions by enhancing the hallucinations in the segment-level descriptions of videos. These diverse hallucination segment descriptions enable video description models to learn to reduce such hallucinations when describing videos. Using this method, hallucinations generated during video description generation can be taken into account. A dataset can be constructed to train the model to reduce or avoid hallucinations during video description generation, thereby providing a data foundation for training the model.
[0110] like Figure 7 The diagram shows a flowchart of a video segmentation method provided in an embodiment of this application. The video segmentation method is applied in an electronic device and includes the following steps.
[0111] S701, based on the scene similarity between adjacent video frames in the video, the video is segmented into scenes to obtain the segmented video corresponding to each scene.
[0112] In one embodiment of this application, scene similarity between adjacent video frames can be determined by performing scene detection on the video frames in the video. For example, scene features of each video frame can be determined based on algorithms such as image feature extraction, color histogram, and edge detection, and scene similarity can be determined based on the similarity between the scene features of adjacent video frames.
[0113] In one embodiment of this application, adjacent video frames with a scene similarity greater than or equal to a preset similarity threshold can be divided into the same segmented video. When a video frame with a scene similarity less than the similarity threshold is detected, video segmentation can be performed at the detected video frame, thereby obtaining multiple segmented videos containing a single scene. The similarity threshold can be set based on actual needs, and this application does not impose specific limitations on it.
[0114] In one example, if the scene similarity between adjacent video frames from the i-th to the (i+k)-th video frames is greater than or equal to the similarity threshold, and the scene similarity between the (i+k)-th video frame and the (i+k+1)-th video frame is less than the similarity threshold, the video can be segmented at the (i+k+1)-th video frame. This ensures that the i-th to (i+k)-th video frames belong to the same segment, while the (i+k+1)-th video frame and the (i+k)-th video frame do not belong to the same segment. Here, i and k represent positive integers.
[0115] In another embodiment, a preset video segmentation tool can be used to segment the video into scenes, resulting in segmented videos corresponding to each scene. For example, the video segmentation tool may include, but is not limited to, Auto Shot.
[0116] For example Figure 8 The diagram shown is an example of video segmentation and video paragraphs provided in an embodiment of this application. Each video segment is illustrated by a screenshot of its first video frame. The number in the upper left corner represents the time stamp (e.g., mm:ss.fff (minutes:seconds.frame value)) corresponding to the first video frame of the segment in the complete video. For example, the upper left corner number 00:04.15 of the first video frame of segment Q2 indicates that it is the 15th frame at minute 0, second 4 of the complete video; the upper left corner number 00:11.23 of the first video frame of segment Q3 indicates that it is the 23rd frame at minute 0, second 11 of the complete video.
[0117] In one embodiment of this application, the duration of each segmented video can be determined. The duration can be expressed in frames or in time units (e.g., seconds, milliseconds, etc.), and this application does not impose any restrictions on this. For example, if the frame rate of the video is 30 frames / second, the first video frame of segmented video Q2 is located at the 15th frame of the 4th second of the 0th minute of the complete video, and the first video frame of segmented video Q3 is located at the 23rd frame of the 11th second of the 0th minute of the complete video, the duration of segmented video Q2 can be determined as (11*30+23)-(4*30+15)=218 frames, or =218 frames / 30≈7.27 seconds.
[0118] Through the above embodiments, a video can be divided into multiple independent segmented videos, so that each segmented video can focus on a specific scene or plot, thereby facilitating subsequent analysis and description.
[0119] S702 splices the segmented video based on preset duration parameters to obtain video segments.
[0120] In one embodiment of this application, the duration of the segmented video may be relatively short. Multiple consecutive segmented videos can be spliced together based on preset time parameters to obtain multiple video segments.
[0121] In one example, the time parameter can represent the target range of the video segment's duration, which can be set based on actual needs. For example, the time parameter could be 1 minute to 1 minute 15 seconds, 1 minute 10 seconds to 1 minute 20 seconds, etc. Alternatively, if the frame rate is 30 frames per second, the time parameter could be 1800 frames to 2250 frames, 2100 frames to 2400 frames, etc.
[0122] In one example, the number of video segments can be determined based on the splicing result, for example, the number of video segments may be 2, 3, 5, etc., and this application does not impose a specific limitation on this.
[0123] refer to Figure 8 As shown, if the time parameter is between 1 minute and 1 minute 15 seconds, segmented videos Q1 to Q6 can be spliced together to obtain video segment 1. Segmented videos Q7 to Q10 can be spliced together to obtain video segment 2. The duration of video segment 1 is approximately 1 minute 12 seconds. Although the duration of video segment 2 may not fall within the range of 1 minute to 1 minute 15 seconds, video segments 1 and 2 already contain the entire content of the complete video, therefore they do not affect subsequent processing.
[0124] Through the above embodiments, multiple consecutive segmented videos are merged into a longer video segment, which not only reduces the number of subsequent video descriptions and thus improves processing efficiency, but also avoids the accumulation of too many hallucinations in the video description, thereby improving the accuracy of the generated video description.
[0125] The hallucination-enhanced data generation method provided in the above embodiments takes into account the hallucinations that may occur during the video description model's generation process. It constructs a dataset specifically for training or optimizing the video description model to reduce or avoid hallucinations, thus providing a data foundation for training the model. Next, we will introduce a method for optimizing the model using the dataset obtained from the hallucination-enhanced data generation method.
[0126] See Figure 9The diagram shown is a flowchart of a model optimization method provided in an embodiment of this application. The model optimization method is applied to electronic devices and includes the following steps.
[0127] S901 obtains multiple segment-level descriptions of multiple video segments through a preset video description model.
[0128] In one embodiment of this application, the specific implementation of step S901 can be referred to the description in steps S601 to S602, and will not be repeated here.
[0129] S902, obtain multiple paragraph-level descriptions, including hallucinations, corresponding to multiple paragraph-level descriptions.
[0130] In one embodiment of this application, the specific implementation of step S902 can be referred to the description in step S603, and will not be repeated here.
[0131] S903 performs supervised optimization of the video description model based on multiple paragraph-level descriptions and multiple paragraph-level descriptions including hallucinations.
[0132] In one embodiment of this application, referring to step S603, the original paragraph-level description and the corresponding paragraph-level description including hallucinations can be used as a first data pair. A supervised learning method is used to train the video description model using the first data pair, so that the model learns to distinguish between the original paragraph-level description and the corresponding paragraph-level description including hallucinations, thereby enabling the trained video description model to have the ability to generate accurate paragraph-level descriptions and avoid hallucinations.
[0133] In one example, supervised optimization of the video description model includes: inputting a video segment and its corresponding first data pair (segment-level description and corresponding segment-level description including hallucinations) into the video description model; determining whether the first data pair contains segment-level descriptions including hallucinations to obtain a first judgment result. If the accuracy of the first judgment result of the video description model is less than a preset accuracy threshold, the model parameters are adjusted and optimized, and the model's performance is gradually improved through iterative training until a video description model with an accuracy greater than or equal to the preset accuracy threshold is obtained.
[0134] Through the above embodiments, a supervised learning method is employed. The original paragraph-level descriptions and corresponding paragraph-level descriptions, including those involving hallucinations, are input as paired training data into the video description model, enabling the model to learn and distinguish between the two. Through iterative training and optimization, the video description model gradually improves its accuracy in identifying hallucination descriptions until a preset threshold is reached. This allows the video description model to generate accurate paragraph-level descriptions and effectively avoid hallucinations.
[0135] See Figure 10The diagram shown is a flowchart of a data generation method based on illusion optimization provided in an embodiment of this application. The method is applied in an electronic device and includes the following steps.
[0136] S1001, aggregate multiple segment-level descriptions containing hallucinations to obtain an initial video-level description.
[0137] In one embodiment of this application, multiple paragraph-level descriptions containing hallucinations can be obtained based on the hallucination-enhanced data generation method provided in the above embodiments. For example, referring to the descriptions in S601 to S603, the multiple paragraph-level descriptions containing hallucinations obtained in S603 can be used in this embodiment. In another example, referring to the descriptions in S601 to S602, the multiple paragraph-level descriptions obtained in S602 can be used as multiple paragraph-level descriptions containing hallucinations and applied to this embodiment. For example, referring to... Figure 12 The flowchart shown is a data generation method based on illusion enhancement and optimization provided in an embodiment of this application.
[0138] In one embodiment of this application, a preset natural language processing model can be used to aggregate multiple paragraph-level descriptions containing hallucinations (hereinafter referred to as multiple paragraph-level descriptions). In one example, the natural language processing model can be a large language model, which has the ability to analyze and process text data and can aggregate multiple paragraph-level descriptions to obtain an initial video-level description. For example, the large language model can include, but is not limited to, Llama3.
[0139] In one embodiment of this application, multiple paragraph-level descriptions and preset second prompt words can be input into a large language model. The large language model analyzes and aggregates the multiple paragraph-level descriptions based on the second prompt words, and outputs an aggregated description of the multiple paragraph-level descriptions as an initial video-level description. The second prompt words can be information (e.g., phrases or sentences) used to guide the model in aggregating the multiple paragraph-level descriptions.
[0140] In one example, the second prompt could be represented as: "You have received some language descriptions about a long video. These descriptions are non-overlapping and accurately cover the entire video. Here is the description: ${segment_narration}. Please give me a ${num_words} word description of the video and return only a summary description. Keep all information complete and do not omit anything. Do not simply merge description entries by scenario." Here, ${segment_narration} represents multiple segment-level descriptions (e.g., segment-level description 1, segment-level description 2, and segment-level description 3 above), and ${num_words} is used to limit the number of words in the initial video-level description.
[0141] Through the above examples, an initial video-level description corresponding to each video can be obtained based on a large language model, thereby providing a data foundation for the subsequent illusion optimization process.
[0142] In one example, aggregating paragraph-level descriptions 1, 2, and 3 yields an initial video-level description 123, which can be exemplarily represented as: "This video is a stunning long-form piece that takes viewers on a journey through the African savanna, showcasing the resilience and majesty of cattle and elephants in their natural habitat. The video begins with a herd of cattle traversing barren land, steadily moving forward, kicking up clouds of dust as they cross the vast, dusty plains. The scene is set under a hazy, overcast sky, emphasizing the arid environment. Close-up shots capture the intricate details of the cattle's red skin and their expressive, large ears, highlighting their majestic presence. The video also shows tender moments between cows and calves, revealing the deep bond and connection between them. The background is set on a vast, open grassland with sparse vegetation, emphasizing the immensity of the African landscape. The video also presents a breathtaking aerial view of the expansive landscape, showcasing the stark contrast between the deep, dark waters of a large lake and the surrounding arid, green terrain. The overall atmosphere is one of vibrant, otherworldly beauty, with natural elements harmoniously blending together to create a captivating visual experience."
[0143] In one embodiment of this application, the initial video description may contain: illusions in each segment-level description, and illusions accumulated or newly generated during the aggregation process of multiple segment-level descriptions. Thus, by optimizing the numerous illusions present in the initial video description in subsequent processes, an optimized video-level description for optimizing the model can be obtained, as detailed in the following embodiments.
[0144] S1002, the initial video-level description is optimized based on the preset second prompt information to obtain an optimized video-level description.
[0145] In one embodiment of this application, the preset second prompt information may include, but is not limited to, one or more of the following: hallucination judgment prompts and hallucination optimization prompts. The hallucination judgment prompts instruct the natural language processing model to generate multiple hallucination judgment questions for the initial video-level description, and provide a judgment result for each hallucination judgment question. For example, based on the hallucination judgment questions, the answer to the question is determined, and then the corresponding judgment result is obtained. Based on the hallucination judgment prompts in the second prompt information, the hallucinations present in the initial video-level description can be determined according to multiple paragraph-level descriptions. The hallucination optimization prompts instruct the natural language model to optimize the hallucinations in the judgment results of the hallucination judgment prompts. Based on the hallucination optimization prompts in the second prompt information, the hallucinations can be optimized to obtain an optimized video-level description.
[0146] In one embodiment of this application, multiple paragraph-level descriptions, an initial video-level description, and second prompt information can be input into a pre-trained natural language processing (NLP) model. The NLP model generates multiple illusion judgment questions based on illusion judgment prompts for the initial video-level description. The answers to the multiple illusion judgment questions are determined, and judgment results are determined based on the answers. Illusions in the initial video-level description are then identified based on the judgment results. The NLP model optimizes the prompts based on illusions to eliminate the identified illusions, resulting in a corresponding optimized video-level description. For example, the NLP model may include, but is not limited to, Llama3 (Large Language Model Architecture 3.0).
[0147] In one embodiment of this application, each illusion judgment question can correspond to a type of illusion generation prompt. For example, in conjunction with the above description of S603, the types of illusion generation prompts include one or more of the following: logical illusion generation prompts, redundant illusion generation prompts, entity illusion generation prompts, attribute illusion generation prompts, quantity illusion generation prompts, and behavioral illusion generation prompts. Since each illusion judgment question can correspond to a type of illusion generation prompt, illusions in the initial video-level description can be determined based on illusion judgment questions across multiple dimensions, avoiding omissions.
[0148] In one example, a hallucination judgment prompt could be represented as: "I have a summary description of a video, which is derived from multiple segment descriptions arranged chronologically within the video. Some parts of the summary description may contain hallucinations, omissions, or repetitions, so I want to use a VQA model to carefully examine some key descriptions. Please ask up to {question_num} of the most important and specific questions so that I can double-check and improve the accuracy of the summary description. These questions should determine whether the details of all objects mentioned in the summary description are consistent with the segment descriptions, including the object's name, attributes, quantity, and actions (or events). These questions should also determine whether there is semantic redundancy and logical errors in the summary description. Summary description: {summary_narration}. Segment description: {segment_narration}. Please output the questions and corresponding answers regarding the differences between the summary description and the segment descriptions."
[0149] Here, VQA stands for Visual Question Answering. {question_num} is used to limit the number of illusion judgment questions generated, summary_narration represents the initial video-level description (e.g., initial video-level description 123), and segment_narration represents multiple segment-level descriptions (e.g., segment-level description 1, segment-level description 2, and segment-level description 3).
[0150] In one example, the natural language model generates a hallucination judgment question 1 based on the aforementioned hallucination judgment prompt, which can be exemplarily represented as: "What type of animal appears in the video?" Based on hallucination judgment question 1, the natural language processing model identifies the animal type in the summary description fragment and obtains the question answers: "Summary: Cow and elephant" and "Fragment: Elephant (no mention of cow)". Based on the question answers, the judgment result is determined as: "The summary description is incorrect because the fragment description only mentions elephant and does not mention cow." Here, "incorrect" indicates the existence of a hallucination.
[0151] In one example, the natural language model generates a hallucination judgment question 2 based on the aforementioned hallucination judgment prompt, which can be exemplarily represented as: "Are there close-up shots of animals in the video?" Based on hallucination judgment question 2, the natural language processing model identifies close-up shots of animals in the summary description and the segment description, obtaining the question answers: "Summary: Yes, there are close-up shots of the cow's red skin and big ears," and "Segment: Yes, there are close-up shots of the elephant's big ears and skin details." Based on the question answers, the judgment result is determined as: "The summary description is incorrect because it mentions close-up shots of cows, but the segment description only mentions elephants."
[0152] In one example, the natural language model generates illusion judgment question 3 based on the aforementioned illusion judgment prompt, which can be exemplarily represented as: "Are there any heartwarming moments between animals in the video?" Based on illusion judgment question 3, the natural language processing model identifies heartwarming moments between animals in the summary description and the segment description, obtaining the question answers: "Summary: Yes, between a cow and a calf," and "Segment: Yes, between a mother elephant and a calf." Based on the question answers, the judgment result is determined as: "The summary description is incorrect because it mentions a cow and a calf, but the segment description only mentions a mother elephant and a calf."
[0153] In one example, the natural language model generates illusion judgment question 4 based on the aforementioned illusion judgment prompt, which can be exemplarily represented as: "Is there any repeated description in the summary?" Based on illusion judgment question 4, the natural language processing model identifies whether there is repeated description in the summary description and the fragment description, obtaining the question answers: "Summary: The video shows a herd of cattle moving steadily through a desolate landscape, kicking up clouds of dust. (This sentence is repeated with slight variations)", "Fragment: There is no repeated description". Based on the question answers, the judgment result is determined as: "The summary description contains repeated descriptions, which may indicate a lack of clarity or accuracy."
[0154] In one example, the natural language model generates a hallucination judgment question 5 based on the aforementioned hallucination judgment prompt, which can be exemplarily represented as: "Is there a logical error in the summary?" Based on hallucination judgment question 5, the natural language processing model identifies whether there are logical errors in the summary description and the fragment description, obtaining the question answer: "At the beginning of the video, a herd of cows crosses a desolate landscape... (this implies that the video begins with cows, but the fragment description only mentions elephants)", "Fragment: No logical error". Based on the question answer, the judgment result is determined as: "The summary description contains a logical error because it implies that the video begins with cows, which is not supported by the fragment description."
[0155] In one example, the natural language model generates illusion judgment question 6 based on the aforementioned illusion judgment prompt, which can be exemplarily represented as: "Is there an inconsistency in the description of the environment in the summary?" Based on illusion judgment question 6, the natural language processing model identifies whether there are logical errors in the summary description and the fragment description, obtaining the question answer: "Summary: The scene is set under a hazy sky, emphasizing the arid environment. The video also shows a stunning aerial view of the vast landscape, demonstrating the stark contrast between the deep, dark waters of the Great Lake and the surrounding arid, green terrain," and "Fragment: The environment is consistently described as arid, with no mention of green terrain." Based on the question answer, the judgment result is determined as: "The summary description is inconsistent because it describes the environment as arid but also mentions green terrain, which is not supported by the fragment description."
[0156] Through the above embodiments, various hallucination judgment questions can be generated based on hallucination judgment prompts. Based on the hallucination judgment questions and their answers, corresponding judgment results can be obtained, thereby determining the hallucination in the initial video-level description from multiple dimensions, thus providing a foundation for subsequent hallucination optimization.
[0157] In one embodiment of this application, the hallucination optimization prompt may include an optimization reason prompt. Based on the optimization reason prompt, a natural language model can generate reasons for optimizing the hallucination. For example, the reasons for optimizing the hallucination may include the type of hallucination and the cause of its occurrence. The natural language model can then make targeted modifications to the hallucination based on the reasons for optimization. In another example, the hallucination optimization prompt may also include an optimization content prompt. Based on the optimization content prompt, the natural language model can obtain the optimization content for the hallucination.
[0158] In one example, the hallucination optimization hint 1 could be represented as: "I have a summary description of a video. Some parts of the summary description may contain hallucinations, omissions, or repetitions, so I used a VQA model to carefully examine some key descriptions. Here is the original summary description that may contain hallucinations: {summary_narration}. Here is the question-and-answer content from the VQA model: {QA}. Please correct the summary description if it is inconsistent with the VQA. This is a difficult question. Please minimize the modifications to the summary description and list all changes and their reasons. Please give the output in the following format: Reason: Changes made: Final refined description:".
[0159] Here, `summary_narration` represents the initial video-level description (e.g., initial video-level description 123), and the VQA model's question-and-answer content: `{QA}` represents the judgment result of the hallucination judgment prompt. "List all changes and their reasons" indicates the reason for optimization, and "Changes made" indicates the optimized content. "Final refined description" is used to instruct the model to output an optimized video-level description.
[0160] In one example, based on the aforementioned illusion optimization tip 1, the optimization rationale output by the natural language model can be represented as follows: "Reason: The summary description contains illusions, omissions, and repetitions, which may lead to inconsistencies and inaccuracies. The VQA model identified several issues, including incorrect mentions of cows, omissions of lakes and bird's-eye views, and inconsistencies in the environmental description."
[0161] In one example, based on the aforementioned illusion optimization tip 1, the optimized content output by the natural language model can be represented as follows: "Changes made: Removed mention of cows: The snippet description only mentions elephants, so including cows in the summary description is incorrect. Added lakes and bird's-eye view: The snippet description mentions lakes and bird's-eye view, which are important features of the video. Adding these details will improve the accuracy and completeness of the summary description. Corrected environment description: The snippet description consistently describes the environment as arid, without mentioning green terrain. Changing the summary description to be consistent with this will eliminate the inconsistency."
[0162] In one example, based on the aforementioned illusion optimization tip 1, the optimized video-level description output by the natural language model can be represented as follows: "This video is a stunning long-form piece that takes viewers on a journey through the African savanna, showcasing the resilience and majesty of elephants in their natural habitat. The video begins with a herd of elephants steadily moving across a desolate landscape, kicking up clouds of dust as they traverse the vast, dusty plains. The scene is set under a gloomy sky, highlighting the arid environment. Close-up shots capture the intricate details of the elephants' skin and their large ears, emphasizing their majestic presence. The video also shows heartwarming moments between mother and calf elephants, demonstrating the deep bond and connection between them. The background is set on a vast, open grassland with sparse vegetation, emphasizing the immensity of the African landscape. The video also presents a breathtaking aerial view of the expansive landscape, showcasing the stark contrast between the deep, dark waters of the Great Lake and the surrounding arid terrain. The overall atmosphere is one of vibrant, otherworldly beauty, with natural elements harmoniously blending together to create a captivating visual experience."
[0163] Through the above embodiments, the natural language model optimizes the hallucinations in the initial video-level description based on hallucination optimization prompts. The number of hallucinations in the final optimized video-level description is significantly reduced compared to the initial video-level description. Therefore, the optimized video-level description and the corresponding optimized video-level description can be used as data pairs to train the natural language model.
[0164] In one embodiment of this application, the initial video-level description and the corresponding optimized video-level description can be used as a pair of training data. A supervised learning method is used to train the natural language processing model with the paired training data, so that the model learns to distinguish between the initial video-level description and the corresponding optimized video-level description. This enables the trained natural language processing model to have the ability to aggregate paragraph-level descriptions into accurate video-level descriptions and avoid illusions.
[0165] In one example, the optimized video-level description does not exhibit hallucinations. In another example, even if the optimized video-level description does exhibit hallucinations, the number of hallucinations is significantly reduced compared to the corresponding initial video-level description. Therefore, using optimized video-level descriptions paired with their corresponding optimized video-level descriptions as data pairs allows for model training.
[0166] The hallucination-optimized data generation method provided in this application aggregates paragraph-level descriptions into initial video-level descriptions using a natural language processing (NLP) model. The initial video-level descriptions are then optimized to reduce video-level hallucinations. The resulting optimized video-level descriptions can be used to train the NLP model, thereby reducing the accumulation of hallucinations during the aggregation of paragraph-level descriptions. Using this method, hallucinations during video description aggregation can be considered, and a dataset can be constructed to train the model to reduce or avoid hallucinations during the aggregation process, thus providing a data foundation for model training.
[0167] Next, a method for optimizing the model using a dataset obtained from a data generation method based on illusion optimization will be introduced. In one embodiment of this application, based on reference... Figure 11 The diagram shown is a flowchart of a model optimization method provided in another embodiment of this application. The model optimization method is applied to electronic devices and includes the following steps.
[0168] S1101 aggregates multiple paragraph-level descriptions using a pre-defined natural language processing model to obtain an initial video-level description.
[0169] In one embodiment of this application, the specific implementation of step S1101 can be referred to the description in step S1001, and will not be repeated here.
[0170] S1102, Obtain the optimized video-level description corresponding to the initial video-level description.
[0171] In one embodiment of this application, the specific implementation of step S1102 can be referred to the description in step S1002, and will not be repeated here.
[0172] S1103 performs supervised optimization of the natural language processing model based on the initial video-level description and the optimized video-level description.
[0173] In one embodiment of this application, referring to step S1002, the initial video-level description and the corresponding optimized video-level description can be used as a second data pair. A supervised learning method is used to train the video description model using the second data pair, so that the model learns to distinguish between the initial video-level description and the corresponding optimized video-level description. This enables the trained video description model to have the ability to aggregate segment-level descriptions into accurate video-level descriptions and avoid hallucinations.
[0174] In one example, supervised optimization of the natural language processing (NLP) model includes: inputting a video and its corresponding second data pair (an initial video-level description and a corresponding optimized video-level description) into the NLP model; using the NLP model to judge the optimized video-level description in the second data pair to obtain a second judgment result; if the accuracy of the second judgment result of the NLP model is less than a preset accuracy threshold, adjusting and optimizing the model's parameters; and gradually improving the model's performance through iterative training until an NLP model with an accuracy greater than or equal to the preset accuracy threshold is obtained.
[0175] Through the above embodiments, a supervised learning method is employed, inputting initial video-level descriptions and corresponding optimized video-level descriptions as paired training data into the natural language processing model. This enables the model to learn and distinguish between the two. Through iterative training and optimization, the natural language processing model gradually improves its accuracy in judging hallucination descriptions until it reaches a preset threshold. This allows the natural language processing model to aggregate paragraph-level descriptions into accurate video-level descriptions and effectively avoid hallucinations.
[0176] In another embodiment of this application, supervised optimization of the video description model can be performed based on the initial video-level description and the optimized video-level description. In one example, a video and its corresponding second data pair (the initial video-level description and the corresponding optimized video-level description) are input into the video description model. The video description model then determines the optimized video-level description in the second data pair to obtain a third determination result. If the accuracy of the third determination result of the video description model is less than a preset accuracy threshold, the model's parameters are adjusted and optimized. The model's performance is gradually improved through iterative training until a natural language processing model with an accuracy greater than or equal to the preset accuracy threshold is obtained.
[0177] Through the above embodiments, a supervised learning method is employed, inputting the initial video-level description and its corresponding optimized video-level description as paired training data into the video description model. This enables the model to learn and distinguish between the two. Through iterative training and optimization, the video description model gradually improves its accuracy in judging hallucination descriptions until it reaches a preset threshold, thereby enabling the video description model to generate accurate video-level descriptions and effectively avoid hallucinations.
[0178] The model optimization method provided in this application embodiment can perform supervised optimization of video description models and natural language processing models based on training data generated by data generation methods. This effectively improves the accuracy of video description models and natural language processing models in judging hallucination descriptions, thereby enabling the optimized video description models and natural language processing models to reduce the possibility of hallucinations in the description text and generate more accurate video description text.
[0179] refer to Figure 12The diagram shows a flowchart of a data generation method based on hallucination enhancement and optimization provided in an embodiment of this application. In this embodiment, hallucination enhancement can be performed on paragraph-level descriptions to generate diverse hallucination paragraph descriptions; the paragraph-level descriptions can be aggregated, and the aggregated data can be further optimized. This allows for the construction of a dataset specifically designed to reduce or avoid hallucinations during the generation and aggregation of video descriptions, thus providing a data foundation for training the model. This method, when applied to electronic devices, includes the following steps.
[0180] S1201, the video is segmented to obtain multiple video segments.
[0181] In one embodiment of this application, the specific implementation of step S1201 can be referred to the description in step S601, and will not be repeated here.
[0182] S1202, obtain the paragraph-level description corresponding to each video segment.
[0183] In one embodiment of this application, the specific implementation of step S1202 can be referred to the description in step S602, and will not be repeated here.
[0184] S1203, based on the preset first prompt information, update the data of multiple paragraph-level descriptions to obtain multiple paragraph-level descriptions including hallucinations.
[0185] In one embodiment of this application, the specific implementation of step S1203 can be referred to the description in step S603, and will not be repeated here.
[0186] S1204, aggregate the paragraph-level descriptions to obtain the initial video-level descriptions.
[0187] In one embodiment of this application, the segment-level description in step S1204 may be the segment-level description corresponding to each video segment obtained in step S1202. Specific embodiments for aggregating segment-level descriptions can be found in the description in step S1001, and will not be repeated here.
[0188] S1205, the initial video-level description is optimized based on the preset second prompt information to obtain an optimized video-level description.
[0189] In one embodiment of this application, the specific implementation of step S1205 can be referred to the description in step S1002, and will not be repeated here.
[0190] In one example, the video description model can be optimized using paragraph-level descriptions and paragraph-level descriptions including hallucinations. For example, refer to the relevant embodiments in steps S901 to S903, which will not be repeated here. The natural language model can be optimized using initial video-level descriptions and optimized video-level descriptions. For example, refer to the relevant embodiments in steps S1101 to S1103, which will not be repeated here.
[0191] The data generation method based on illusion enhancement and optimization provided in this application enhances the illusion of segment-level descriptions in videos to generate diverse illusion segment descriptions. These diverse illusion segment descriptions enable video description models to learn to reduce such illusions when describing videos. A natural language processing (NLP) model aggregates segment-level descriptions to form initial video-level descriptions. These initial video-level descriptions are then optimized to reduce video-level illusions. The resulting optimized video-level descriptions can be used to train the NLP model, thereby reducing the accumulation of illusions during the aggregation of segment-level descriptions. Using this method, illusions generated during video description generation and aggregation can be considered. A dataset is constructed to address these illusions and train the model to reduce or avoid illusions during these processes, thus providing a data foundation for model training.
[0192] See Figure 13 The diagram illustrates the functional modules of a data generation method provided in an embodiment of this application. The data generation method includes a video processing module, a segment-level hallucination generation module, and a video-level hallucination optimization module. The video processing module uses the Auto Shot tool to segment the original long video into single-scene video fragments. It then uses a duration hyperparameter to combine the single-scene video fragments into multi-scene video segments. The segment-level hallucination generation module uses the InternVL2 model to generate segment descriptions for the video segments and uses these as initial segment-level descriptions. It constructs hallucination generation cues based on multiple hallucination cues and uses the Llama3 model and hallucination cues to generate hallucinations in the segment-level descriptions, resulting in hallucination-generated segment-level descriptions. The video-level hallucination optimization module uses the Llama3 model to aggregate segment descriptions, resulting in initial video-level descriptions. It constructs hallucination judgment problems based on multiple hallucination judgment problems and uses Llama3 and the hallucination judgment problems to eliminate hallucinations in the initial video-level descriptions, resulting in hallucination-optimized video-level descriptions.
[0193] Through the above embodiments, the video processing module uses the Auto Shot tool to segment long videos into single-scene video fragments and combines them into multi-scene video segments based on the duration hyperparameter. The segment-level hallucination generation module uses the InternVL2 model to generate initial segment descriptions and uses the Llama3 model combined with various hallucination cues to generate segment-level descriptions containing hallucinations. The video-level hallucination optimization module aggregates segment descriptions using the Llama3 model to obtain initial video-level descriptions, and then eliminates hallucinations based on hallucination judgment questions, thus obtaining optimized video-level descriptions. This method can efficiently generate hallucination-enhanced segment-level descriptions and hallucination-eliminating optimized video descriptions. Therefore, the hallucination-enhanced segment-level descriptions and hallucination-eliminating optimized video descriptions can be used to train and optimize the model, improving the model's video description capabilities and avoiding hallucination generation.
[0194] This application also provides an electronic device 100, see reference. Figure 14 As shown, the electronic device 100 may be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, as well as cellular phone, personal digital assistant (PDA), augmented reality (AR) device, virtual reality (VR) device, artificial intelligence (AI) device, wearable device, in-vehicle device, smart home device and / or smart city device. The specific type of electronic device 100 is not specifically limited in the embodiments of this application.
[0195] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, Universal Serial Bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and Subscriber Identification Module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0196] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0197] Processor 110 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.
[0198] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.
[0199] The processor 110 may also include a memory for storing instructions and data. In one embodiment of this application, the memory in the processor 110 is a cache memory. The memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instructions or data again, it can directly retrieve them from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0200] In one embodiment of this application, the processor 110 may include one or more interfaces. These interfaces may include an Inter-integrated Circuit (I2C) interface, an Inter-integrated Circuit Sound (I2S) interface, a Pulse Code Modulation (PCM) interface, a Universal Asynchronous Receiver / Transmitter (UART) interface, a Mobile Industry Processor Interface (MIPI) interface, a General-Purpose Input / Output (GPIO) interface, a Subscriber Identity Module (SIM) interface, and / or a Universal Serial Bus (USB) interface, etc.
[0201] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In one embodiment of this application, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 100.
[0202] The I2S interface can be used for audio communication. In one embodiment of this application, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to realize communication between the processor 110 and the audio module 170. In one embodiment of this application, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to realize the function of answering phone calls through a Bluetooth headset.
[0203] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In one embodiment of this application, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In another embodiment of this application, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0204] The UART interface is a universal serial data bus used for asynchronous communication. The bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In one embodiment of this application, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In one embodiment of this application, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music playback via Bluetooth headphones.
[0205] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a Camera Serial Interface (CSI) and a Display Serial Interface (DSI). In one embodiment of this application, the processor 110 and the camera 193 communicate via the CSI interface to realize the shooting function of the electronic device 100. The processor 110 and the display screen 194 communicate via the DSI interface to realize the display function of the electronic device 100.
[0206] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In one embodiment of this application, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0207] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. Furthermore, the interface can be used to connect other electronic devices 100, such as AR devices.
[0208] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0209] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via a USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device 100 via the power management module 141.
[0210] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0211] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0212] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.
[0213] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low-noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In one embodiment of this application, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In another embodiment of this application, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.
[0214] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through audio devices (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In one embodiment of this application, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and housed within the same device as the mobile communication module 150 or other functional modules.
[0215] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including Wireless Local Area Networks (WLANs) (such as Wireless Fidelity (Wi-Fi) networks), Bluetooth (BT), Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), and Infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0216] In one embodiment of this application, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the Beidou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).
[0217] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0218] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a Liquid Crystal Display (LCD), an Organic Light-Emitting Diode (OLED), an Active-Matrix Organic Light-Emitting Diode (AMOLED), a Flexible Light-Emitting Diode (FLED), a Minied, Microled, Micro-OLED, or a Quantum Dot Light-Emitting Diode (QLED), etc. In one embodiment of this application, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.
[0219] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0220] The ISP is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, converting it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In one embodiment of this application, the ISP can be set in the camera 193.
[0221] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In one embodiment of this application, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0222] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.
[0223] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record video in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0224] NPU stands for Neural Network (NN) computing processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs can enable intelligent cognitive applications in electronic devices, such as image recognition, data generation, speech recognition, and text understanding.
[0225] Internal memory 121 may include one or more random access memory (RAM) and one or more non-volatile memory (NVM).
[0226] Random access memory can include static random-access memory (SRAM), dynamic random-access memory (DRAM), synchronous dynamic random-access memory (SDRAM), and double data rate synchronous dynamic random-access memory (DDR SDRAM, such as fifth-generation DDR SDRAM, which is generally called DDR5 SDRAM).
[0227] Non-volatile memory can include disk storage devices and flash memory.
[0228] Flash memory can be classified according to its operating principle, including NOR FLASH, NAND FLASH, 3D NAND FLASH, etc.; according to the level of the storage cell, including single-level cell (SLC), multi-level cell (MLC), triple-level cell (TLC), quad-level cell (QLC), etc.; and according to the storage specification, including universal flash storage (UFS) and embedded multi-media card (eMMC), etc.
[0229] The random access memory can be directly read and written by the processor 110. It can be used to store executable programs (such as machine instructions) of the operating system or other running programs, as well as user and application data.
[0230] Non-volatile memory can also store executable programs and user and application data, and can be pre-loaded into random access memory for direct reading and writing by the processor 110.
[0231] The external memory interface 120 can be used to connect to external non-volatile memory, thereby expanding the storage capacity of the electronic device 100. The external non-volatile memory communicates with the processor 110 through the external memory interface 120 to perform data storage functions. For example, music, video, and other files can be stored in the external non-volatile memory.
[0232] Internal memory 121 or external memory interface 120 is used to store one or more computer programs. The one or more computer programs are configured to be executed by processor 110. The one or more computer programs include multiple instructions, which, when executed by processor 110, can implement the screen display detection method executed on electronic device 100 in the above embodiments, so as to realize the screen display detection function of electronic device 100.
[0233] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0234] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In one embodiment of this application, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0235] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.
[0236] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.
[0237] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0238] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.
[0239] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.
[0240] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.
[0241] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.
[0242] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In one embodiment of this application, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100. This application also provides a computer storage medium storing computer instructions. When the computer instructions are executed on the electronic device 100, the electronic device 100 performs the aforementioned related method steps to implement the data generation method and model optimization method in the above embodiments.
[0243] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement the data generation method and model optimization method described in the above embodiments.
[0244] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component or module. The apparatus may include a connected processor and a memory; wherein, the memory is used to store computer execution instructions, and when the apparatus is running, the processor may execute the computer execution instructions stored in the memory to cause the chip to execute the data generation method and model optimization method in the above method embodiments.
[0245] In this embodiment, the electronic device, computer storage medium, computer program product or chip are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding method provided above, and will not be repeated here.
[0246] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0247] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0248] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0249] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0250] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0251] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of this application without departing from the spirit and scope of the technical solutions of this application.
Claims
1. A data generation method based on illusion enhancement, applied to electronic devices, characterized in that, The method includes: The video is segmented into multiple video segments; Use a video description model to obtain the paragraph-level description corresponding to each video segment; Based on multiple illusion generation prompts in the preset first prompt information, at least one data update is performed on each paragraph-level description, so that each updated paragraph-level description contains at least one illusion corresponding to an illusion generation prompt, resulting in multiple paragraph-level descriptions including illusions; wherein, each paragraph-level description with a corresponding relationship and the paragraph-level descriptions including illusions are used to train the video description model; the multiple illusion generation prompts include at least one of logical illusion generation prompts and redundant illusion generation prompts, as well as at least one of entity illusion generation prompts, attribute illusion generation prompts, quantity illusion generation prompts, and behavior illusion generation prompts; The logical illusion generation prompt is used to instruct the text order of each paragraph-level description to be randomly shuffled; The redundant illusion generation prompt is used to instruct the random repetition of the text statements described at each paragraph level; The entity illusion generation prompt is used to instruct the entity objects of each paragraph-level description to be updated; The attribute illusion generation prompt is used to instruct the adjectives of each paragraph-level description to be updated; The quantity illusion generation prompt is used to instruct the updating of the quantity words in each paragraph-level description; The behavioral illusion generation prompt is used to instruct the verbs in each paragraph-level description to be updated.
2. The data generation method based on hallucination enhancement as described in claim 1, characterized in that, The process of segmenting the video to obtain multiple video segments includes: Based on the scene similarity between adjacent video frames in the video, the video is segmented into scenes to obtain segmented videos corresponding to each scene; The segmented video is spliced together based on preset duration parameters to obtain the video segment.
3. The data generation method based on hallucination enhancement as described in claim 1, characterized in that, The multiple hallucination generation prompts include a first hallucination generation prompt and / or a second hallucination generation prompt, wherein the first hallucination generation prompt includes prompt information for causing the language expression of each paragraph-level description to generate hallucination, and the second hallucination generation prompt includes prompt information for causing the entity object of each paragraph-level description to generate hallucination.
4. A data generation method based on illusion optimization, applied to electronic devices, characterized in that, The method includes: A natural language description model is used to aggregate multiple paragraph-level descriptions containing hallucinations to obtain an initial video-level description. The multiple paragraph-level descriptions containing hallucinations are obtained using the hallucination-enhanced data generation method as described in any one of claims 1 to 3. The initial video-level description is optimized based on the preset second prompt information to obtain an optimized video-level description, including: based on the multiple paragraph-level descriptions containing hallucinations, and based on the second prompt information, identifying and removing hallucinations in the initial video-level description, wherein the hallucinations include hallucinations generated by aggregating the paragraph-level descriptions; The initial video-level description and the optimized video-level description, which have a corresponding relationship, are used to train the natural language description model.
5. The data generation method based on illusion optimization as described in claim 4, characterized in that, The optimization of the initial video-level description based on preset second prompt information to obtain an optimized video-level description includes: Based on the hallucination judgment prompt in the second prompt information, the hallucination present in the initial video-level description is determined according to the multiple paragraph-level descriptions containing hallucinations; Based on the illusion optimization prompt in the second prompt information, the illusion is optimized to obtain the optimized video-level description.
6. The data generation method based on illusion optimization as described in claim 5, characterized in that, The illusion optimization prompt includes a prompt explaining the reasons for optimization, and the method further includes: Based on the optimization reasoning suggestions, a reason for optimizing the illusion is generated.
7. The data generation method as described in claim 6, characterized in that, The hallucination optimization prompt includes optimization content prompts, and the method further includes: Based on the aforementioned optimized content prompts, the optimized content for the hallucination is obtained.
8. A model optimization method applied to electronic devices, characterized in that, The method includes: Based on video segments, corresponding segment-level descriptions are obtained through a preset video description model; Obtain the paragraph-level description including hallucinations corresponding to the paragraph-level description, wherein the paragraph-level description including hallucinations is obtained according to the data generation method based on hallucination enhancement as described in any one of claims 1 to 3; Based on the paragraph-level description and the paragraph-level description including hallucinations, the video description model is subjected to supervised optimization.
9. A model optimization method applied to electronic devices, characterized in that, The method includes: The paragraph-level descriptions are aggregated using a pre-defined natural language processing model to obtain the initial video-level descriptions. Obtain the optimized video-level description corresponding to the initial video-level description, wherein the optimized video-level description is obtained according to the data generation method based on illusion optimization as described in any one of claims 4 to 7; Based on the initial video-level description and the optimized video-level description, the natural language processing model is subjected to supervised optimization.
10. An electronic device, characterized in that, The electronic device includes a memory and a processor: The memory is used to store program instructions; The processor is configured to read and execute the program instructions stored in the memory, and when the program instructions are executed by the processor, cause the electronic device to perform the data generation method as described in any one of claims 1 to 7, or the model optimization method as described in any one of claims 8 to 9.
11. A computer storage medium, characterized in that, The computer storage medium stores program instructions that, when executed on an electronic device, cause the processor of the electronic device to perform the data generation method as described in any one of claims 1 to 7, or the model optimization method as described in any one of claims 8 to 9.
Citation Information
Patent Citations
Video editing method and device, computing equipment and storage medium
CN114302174A
Intelligent video editing method and system based on large model
CN117812386A
Big language model illusion relieving scheme based on citation correction
CN117910449A
Sample data construction method and question and answer model training method
CN118227731A