Retrieval enhancement cue word optimization method for text-to-video generation

By constructing a relationship diagram and a dual-branch optimization strategy, the problems of insufficient prompt information and inconsistent distribution in the text-generated video model are solved. The generated videos have been significantly improved in content richness and action coherence, achieving higher quality video generation.

CN120354844APending Publication Date: 2025-07-22SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510439689.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, the text-generated video model has insufficient prompt information input to the user, resulting in insufficient motion fluency and timing stability of the generated video, and the prompt is inconsistent with the distribution of training data, which is prone to introduce redundancy or error information.

Method used

The pre-trained sentence transformation model extracts brief prompt features, uses the relationship diagram to retrieve the modifier collection, and uses a two-branch optimization strategy and a fine-tuning discriminant model to generate candidate prompts that match the distribution of the training data.

Benefits of technology

It significantly improves the richness of the text-generated video model content, picture details and action coherence, ensures the static quality and dynamic coherence of the generated video, and improves the stability and robustness of the generation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354844A_ABST
    Figure CN120354844A_ABST
Patent Text Reader

Abstract

The invention provides a retrieval enhancement cue word optimization method for text-to-video generation. The retrieval enhancement cue word optimization method comprises the following steps: extracting features of input brief prompts through a pre-trained sentence conversion model; performing cosine similarity calculation on the features of the brief prompt and the relational graph, and retrieving a modifier set most related to the brief prompt from the relational graph according to a calculation result; processing the brief prompts and the modifier set by adopting a double-branch optimization strategy to obtain candidate prompts; and comparing and selecting the candidate prompts generated by the two branches through a fine-tuning discrimination model to obtain an optimal prompt. According to the method, the problem that brief prompt information is insufficient is solved, and the content richness, picture details and action coherence of the generated video can be remarkably improved through optimized prompt.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a method for optimizing retrieval-enhanced prompts for text-to-video generation. Background Art

[0002] With the success of large-scale text-to-image generation models and text-to-video generation models, recent research has attempted to further improve the performance of text-to-image / video generation models by improving the input text of users. Text-to-image and text-to-video generation models are sensitive to input prompts. However, well-performing prompts are usually model-specific, consistent with training prompts, and inconsistent with user inputs. There are some methods in the prior art that have deeply explored the generation potential of text-to-image and text-to-video generation. These methods mainly focus on prompt optimization for text-to-image models and lack the extension to text-to-video models.

[0003] The prior art has the following defects and deficiencies:

[0004] (1) User prompts are brief and lack information. The prompts provided by users are often too simple and lack the detailed descriptions required for generating high-quality videos. Although some methods have achieved good results in the field of text-to-image generation, they do not fully consider the dynamic characteristics such as temporal continuity, action coherence, and multi-object interaction in video generation, resulting in the generated results being difficult to meet the expected effects in both static picture quality and dynamic coherence, and insufficient optimization of video dynamic characteristics.

[0005] (2) The prompts are inconsistent with the training data distribution. On the one hand, there is a lack of fine guidance. Currently, the commonly used automatic prompt optimization methods mainly use large language models to expand user inputs. However, these methods usually only focus on increasing the vocabulary of descriptions and do not provide detailed guidance on how to select appropriate prompt words and how to construct sentence structures that conform to the training distribution. Although the descriptions are enriched to a certain extent, there are often significant differences from the prompt formats and words used in the training data. On the other hand, it is easy to generate misleading information. Directly using large language models to expand prompts may introduce redundant or even incorrect information. These inaccurate modifications may mislead the generation model, resulting in a deviation between the output video and the user's original intention, thus making the effect of the generated video unstable.

[0006] The defects and deficiencies in the prior art result in significant deficiencies in the action smoothness and temporal stability of the generated videos in the text-to-video generation task, even after the prompts are expanded. Summary of the Invention

[0007] In view of this, the present invention provides a method for optimizing retrieval-enhanced prompts for text-to-video generation to solve the above problems.

[0008] The present invention provides a method for optimizing retrieval-enhanced prompts for text-to-video generation, including: extracting features of an input brief prompt through a pre-trained sentence conversion model; calculating the cosine similarity between the features of the brief prompt and a relationship graph, and retrieving a set of modifiers most relevant to the brief prompt from the relationship graph according to the calculation result; processing the brief prompt and the set of modifiers using a dual-branch optimization strategy to obtain candidate prompts; and comparing and selecting the candidate prompts generated by the above two branches through a fine-tuned discriminative model to obtain the optimal prompt.

[0009] In another implementation manner of the present invention, it further includes: constructing the relationship graph according to the obtained training data, where the relationship graph includes scene, subject, action, and atmosphere descriptions.

[0010] In another implementation manner of the present invention, the processing the brief prompt and the set of modifiers using a dual-branch optimization strategy to obtain candidate prompts includes: calling a frozen large language model through the first branch to fuse the modifiers in the set of modifiers with the brief prompt and perform reconstruction processing on format and grammar to obtain candidate prompts with unified format and conforming to the training data distribution; calling a pre-trained large language model through the second branch to rewrite the brief prompt according to a predefined instruction set to obtain candidate prompts with unified format and conforming to the training data format.

[0011] In another implementation manner of the present invention, the calling a frozen large language model through the first branch to fuse the modifiers in the set of modifiers with the brief prompt and perform reconstruction processing on format and grammar to obtain candidate prompts with unified format and conforming to the training data distribution includes: calling a frozen large language model through the first branch, adopting an iterative fusion strategy, gradually fusing the modifiers in the set of modifiers with the brief prompt to obtain a prompt with enhanced vocabulary; and performing reconstruction on the format and grammar of the prompt with enhanced vocabulary through fine-tuning the large language model to obtain candidate prompts with unified format and conforming to the training data distribution.

[0012] On the other hand, the present invention provides a retrieval-enhanced prompt optimization system for text-to-video generation, including: a feature extraction module: extracting features of an input brief prompt through a pre-trained sentence conversion model; a modifier retrieval module: calculating the cosine similarity between the features of the brief prompt and a relationship graph, and retrieving a set of modifiers most relevant to the brief prompt from the relationship graph according to the calculation result; a branch optimization subsystem: processing the brief prompt and the set of modifiers using a dual-branch optimization strategy to obtain candidate prompts; and a prompt selection module: comparing and selecting the candidate prompts generated by the above two branches through a fine-tuned discriminative model to obtain the optimal prompt.

[0013] On the other hand, the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of a retrieval-enhanced prompt optimization method for text-to-video generation as described in any one of the above are implemented.

[0014] On the other hand, the present invention provides a computer storage medium, characterized in that a computer program is stored on the computer storage medium, and when the computer program is executed by a processor, the steps in a retrieval-enhanced prompt optimization method for text-to-video generation as described in any one of the above are implemented.

[0015] The retrieval-enhanced prompt optimization method for text-to-video generation of the present invention constructs modifier retrieval, sentence reconstruction of prompts, and a prompt selection mechanism through a relational graph. Through the above technical means, the input brief prompt is transformed into a prompt that conforms to the training data distribution, is rich in details, and has a unified format, effectively improving the deficiencies of the current video generation model in dealing with brief prompts, and significantly enhancing the content richness, picture details, and action coherence when the text-to-video generation model generates videos. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. By reading the detailed description of the following embodiments, the advantages and benefits in the solutions will become clear to those skilled in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. In the drawings:

[0017] Figure 1 It is a schematic flowchart of a retrieval-enhanced prompt optimization method for text-to-video generation according to an embodiment of the present invention.

[0018] Figure 2 It is a schematic diagram of the system architecture of a retrieval-enhanced prompt optimization system for text-to-video generation according to an embodiment of the present invention.

[0019] Figure 3 It is a schematic diagram of the construction, composition, and application of a relational graph according to an embodiment of the present invention.

[0020] Figure 4 It is a schematic diagram of the qualitative results of text-to-video generation according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the following will clearly and detailedly describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art shall fall within the scope protected by the embodiments of the present invention.

[0022] Figure 1 The following is a schematic flow chart of a Retrieval-Augmented Prompt Optimization (RAPO) method for text-to-video generation provided by an embodiment of the present invention. As Figure 1 shown, this embodiment mainly includes:

[0023] S101. Extract the features of the input brief prompt through a pre-trained Sentence Transformer.

[0024] S102. Calculate the cosine similarity between the features of the brief prompt and the relationship graph, and retrieve the set of modifiers most relevant to the brief prompt from the relationship graph according to the calculation results.

[0025] S103. Process the brief prompt and the set of modifiers using a double-branch optimization strategy to obtain candidate prompts.

[0026] S104. Compare and select the candidate prompts generated by the above two branches through a fine-tuned discriminant model to obtain the optimal prompt.

[0027] The retrieval-augmented prompt optimization method for text-to-video generation of the present invention transforms the input brief prompt into a prompt that conforms to the training data distribution, is rich in details and has a unified format through a modifier retrieval based on a relationship graph, sentence reconstruction of the prompt, and a prompt selection mechanism. By the above technical means, the deficiencies of the current video generation model in dealing with brief prompts are effectively improved, and the text-to-video (T2V) model can be significantly improved in terms of content richness, picture details, and action coherence when generating videos.

[0028] In another implementation manner of the present invention, it further includes: constructing the relationship graph according to the obtained training data, and the relationship graph includes scene, subject, action, and atmosphere descriptions.

[0029] Exemplarily, as Figure 3As shown, a "relationship graph" is constructed using large-scale training prompts. Each node in the relationship graph represents a scenario or key modifier (such as the subject, action, atmosphere, etc.), and the semantic associations between the words are represented by edges. The modifiers most relevant to the user prompt are extracted from it, and detailed information is automatically added to the concise prompt to make up for the deficiency of the user prompt information.

[0030] In another implementation manner of the present invention, the dual-branch optimization strategy is adopted to process the concise prompt and the modifier set to obtain candidate prompts, including: calling a frozen large language model (LLM) through the first branch to fuse the modifiers in the modifier set with the concise prompt and perform reconstruction processing on the format and grammar, so as to obtain candidate prompts with unified format and conforming to the training data distribution; calling a pre-trained large language model through the second branch to rewrite the concise prompt according to a predefined instruction set, so as to obtain candidate prompts with unified format and conforming to the training data format.

[0031] Exemplarily, the dual-branch optimization strategy includes two parallel optimization branches. One branch directly reconstructs the user prompt using an instruction-tuned language model (LLM), and the other branch expands the prompt based on a lexical enhancement module.

[0032] The first branch uses a frozen large language model (LLM) to iteratively fuse the modifiers, thereby automatically expanding the user's concise prompt into an optimized prompt with rich details and unified format, solving the problems of insufficient prompt information and mismatched training prompt distribution.

[0033] The second branch directly uses a pre-trained large language model to rewrite the user prompt according to a predefined instruction set to generate candidate prompts conforming to the training data format. It adopts a fine-tuned large language model (LLM) with instruction tuning. This model is trained on a large number of rewritten prompt pairs (input prompt and target prompt) to learn how to convert the enhanced prompt into a standard prompt with unified format and conforming to the training data distribution.

[0034] In another implementation of the present invention, the frozen large language model is called through the first branch, and the modifiers in the modifier set are fused with the brief prompt and processed for format and grammar reconstruction to obtain a candidate prompt with a unified format and conforming to the training data distribution, including: calling the frozen large language model through the first branch, adopting an iterative fusion strategy, gradually fusing the modifiers in the modifier set with the brief prompt to obtain a prompt with enhanced vocabulary; and reconstructing the format and grammar of the prompt with enhanced vocabulary through fine-tuning the large language model to obtain a candidate prompt with a unified format and conforming to the training data distribution.

[0035] Exemplarily, the frozen large language model is called to adopt an iterative fusion strategy, gradually fusing each retrieved modifier with the brief prompt, and calling a function (fusion function) at each step to perform a merging operation, thereby generating a prompt with enhanced vocabulary, and using the frozen large language model to ensure adding necessary descriptive details while maintaining the original semantics.

[0036] In another implementation of the present invention, another fine-tuned discriminative large language model is utilized, and its training data includes the original prompt, the prompt after sentence reconstruction, and the candidate prompts directly generated by the pre-trained Rewrite LLM. Through the above-mentioned fine-tuned discriminative model, comprehensive evaluation and selection are performed on the candidate prompts generated by the two branches, and the optimal prompt is selected from the candidate prompts generated by the two branches to ensure that the prompt finally used for video generation is both detailed and conforms to the style of the training data, while taking into account the static picture quality and dynamic coherence, significantly improving the generation effect of the text-to-video (T2V) model.

[0037] Preferably, the present invention does not specifically limit the text-to-video model and the large language model. Based on the solution of the present invention, other text-to-video models can also be used as the base, or other large language models can be used for fine-tuning.

[0038] In summary, the present invention realizes the automatic and detailed optimization of the user input prompt by constructing a relationship graph based on the training data, using the frozen and instruction-tuned large language model to gradually expand and reconstruct the format of the prompt, and selecting the optimal prompt through the discriminant module, providing an input more conforming to the training data style for the text-to-video model, and further improving the static details and dynamic coherence of the generated video.

[0039] The present invention has been verified on the publicly evaluated dataset, achieving significant improvements in static image quality and dynamic video coherence in the text-to-video task, proving the effectiveness of the present invention. The quantitative and qualitative results are shown in Table 1 and Figure 4 as follows.

[0040] Table 1 Quantitative Results on the VBench Dataset

[0041]

[0042] It can be seen that the present invention is superior to other methods in both static dimensions (e.g., visual quality, object category) and dynamic dimensions (e.g., human actions, temporal flicker). Although the optimized prompts of GPT-4 and Open-sora enrich the ordinary prompts provided by users and provide more details describing the scene, object, and action, these long and complex descriptions may confuse the model and even deteriorate the generation effect. It is worth noting that the present invention significantly improves the performance of generating scenes involving more than two subjects.

[0043] The present invention constructs a relational graph using large-scale training data, and realizes the detailed expansion of the brief prompt by retrieving modifiers such as scenes, subjects, actions, and atmospheres related to the user input prompt, making the optimized prompt closer to the distribution of the training data; secondly, a dual-branch design is adopted. One branch iteratively fuses the retrieved modifiers through a frozen large language model, and the other branch directly reconstructs the prompt through a pre-trained model, and automatically selects the optimal candidate through a discriminant model, which not only effectively avoids the misleading information that may be introduced by the pure rewriting method, but also takes into account the semantic preservation and format unity of the prompt; finally, this method shows higher robustness and stability in the static image quality and dynamic coherence of the generated video, especially significantly improving the generation effect in complex scenes and multi-object generation tasks.

[0044] On the other hand of the present invention, as Figure 2 shown, a retrieval-enhanced prompt optimization system for text-to-video generation is provided, including:

[0045] Feature extraction module: Extract the features of the input brief prompt through a pre-trained sentence transformation model.

[0046] Modifier retrieval module: Calculate the cosine similarity between the features of the brief prompt and the relational graph, and retrieve the set of modifiers most relevant to the brief prompt from the relational graph according to the calculation result.

[0047] Branch optimization subsystem: Adopt a dual-branch optimization strategy to process the brief prompt and the set of modifiers to obtain candidate prompts.

[0048] Prompt selection module: Compare and select the candidate prompts generated by the above two branches through a fine-tuned discriminant model to obtain the optimal prompt.

[0049] The retrieval-enhanced prompt optimization system for text-to-video generation of the present invention, through a relational graph to construct modifier retrieval, sentence reconstruction of prompts, and a prompt selection mechanism, transforms the input concise prompt into a prompt that conforms to the training data distribution, is rich in details and has a unified format through the above technical means, effectively improving the deficiencies of the current video generation model in dealing with concise prompts, and enabling the text-to-video generation model to significantly improve in terms of content richness, picture details, and action coherence when generating videos.

[0050] In another implementation manner of the present invention, the branch optimization subsystem includes a vocabulary enhancement module, a sentence reconstruction module, and an instruction rewriting module.

[0051] On the other hand, an electronic device of the present invention includes: a processor, a memory, and a communication bus and a communication interface.

[0052] Wherein:

[0053] The processor, the memory, and the communication interface complete communication with each other through the communication bus.

[0054] The communication interface is used to communicate with other electronic devices or servers.

[0055] The processor is used to execute a program, and specifically can execute the steps of any one of the above-mentioned retrieval-enhanced prompt optimization methods for text-to-video generation in the embodiments.

[0056] Specifically, the program may include program code, and the program code includes computer operation instructions.

[0057] The processor may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0058] The memory is used to store the program. The memory may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0059] The program can specifically be used to cause a processor to execute steps for implementing any of the retrieval-enhanced prompt optimization methods for text-to-video generation described in the embodiments. For the specific implementation of each step in the program, reference can be made to the corresponding descriptions in the steps and units of any of the retrieval-enhanced prompt optimization methods for text-to-video generation in the above steps, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices and modules described above can refer to the corresponding process descriptions in the foregoing method embodiments.

[0060] An exemplary embodiment of the present application also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the methods of the embodiments of the present application.

[0061] The methods according to the embodiments of the present invention described above can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the methods described herein can be stored as such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a RAM, a ROM, a flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods described herein are implemented. In addition, when a general-purpose computer accesses the code for implementing the methods shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0062] So far, specific embodiments of the present invention have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result.

[0063] It should be noted that all directional indications (such as up, down, left, right, back...) in the embodiments of the present invention are only used to explain the relative positional relationship between components in a specific order (as shown in the drawings). If the specific order changes, the directional indications will change accordingly.

[0064] In the description of the present invention, the terms "first" and "second" are only used for conveniently describing different components or names, and cannot be construed as indicating or implying an order relationship, relative importance, or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features.

[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention.

[0066] It should be noted that although the specific embodiments of the present invention have been described in detail with reference to the accompanying drawings, it should not be construed as a limitation on the protection scope of the present invention. Within the scope described in the claims, various modifications and variations that can be made by those skilled in the art without creative efforts still fall within the protection scope of the present invention.

[0067] The examples of the embodiments of the present invention are intended to briefly illustrate the technical features of the embodiments of the present invention, enabling those skilled in the art to intuitively understand the technical features of the embodiments of the present invention, and not serving as an improper limitation on the embodiments of the present invention.

[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A retrieval-enhanced prompt optimization method for text-to-video generation, characterized in that, Including: Extracting the features of the input brief prompt through a pre-trained sentence conversion model; Calculating the cosine similarity between the features of the brief prompt and the relationship graph, and retrieving the set of modifiers most relevant to the brief prompt from the relationship graph according to the calculation result; Processing the brief prompt and the set of modifiers using a dual-branch optimization strategy to obtain candidate prompts; Comparing and selecting the candidate prompts generated by the above two branches through a fine-tuned discriminative model to obtain the optimal prompt.

2. The method according to claim 1, wherein Also including: Constructing the relationship graph according to the obtained training data, where the relationship graph includes scene, subject, action, and atmosphere descriptions.

3. The method according to claim 1, wherein The step of processing the brief prompt and the set of modifiers using a dual-branch optimization strategy to obtain candidate prompts includes: Invoking a frozen large language model through the first branch to fuse the modifiers in the set of modifiers with the brief prompt and perform reconstruction processing on format and grammar to obtain candidate prompts with unified format and conforming to the training data distribution; Invoking a pre-trained large language model through the second branch to rewrite the brief prompt according to a predefined instruction set to obtain candidate prompts with unified format and conforming to the training data format.

4. The method according to claim 3, wherein The step of invoking a frozen large language model through the first branch to fuse the modifiers in the set of modifiers with the brief prompt and perform reconstruction processing on format and grammar to obtain candidate prompts with unified format and conforming to the training data distribution includes: Invoking a frozen large language model through the first branch and adopting an iterative fusion strategy to gradually fuse the modifiers in the set of modifiers with the brief prompt to obtain a prompt with enhanced vocabulary; Performing reconstruction on the format and grammar of the prompt with enhanced vocabulary through fine-tuning the large language model to obtain candidate prompts with unified format and conforming to the training data distribution.

5. A retrieval-enhanced prompt optimization system for text-to-video generation, characterized in that, Including: Feature extraction module: Extracting the features of the input brief prompt through a pre-trained sentence conversion model; Modifier retrieval module: Calculating the cosine similarity between the features of the brief prompt and the relationship graph, and retrieving the set of modifiers most relevant to the brief prompt from the relationship graph according to the calculation result; Branch optimization subsystem: Processing the brief prompt and the set of modifiers using a dual-branch optimization strategy to obtain candidate prompts; Prompt selection module: Comparing and selecting the candidate prompts generated by the above two branches through a fine-tuned discriminative model to obtain the optimal prompt.

6. An electronic device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, it implements the steps of a method for optimizing retrieval-enhanced prompts for text-to-video generation according to any one of claims 1 to 4.

7. A computer storage medium, characterized in that, A computer program is stored on the computer storage medium, and when the computer program is executed by the processor, it implements the steps in a method for optimizing retrieval-enhanced prompts for text-to-video generation according to any one of claims 1 to 4.