Small sample class incremental video action recognition method and device
Through pre-trained contrasting language-visual model and knowledge distillation technology, video frame features are extracted and text features are fused, the problem of video action recognition under small sample conditions is solved, the accuracy of new category recognition is improved, and the forgetting of old categories is reduced, and efficient video action recognition is achieved.
Patent Information
- Application Number
- CN202511049383.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-07-29
AI Technical Summary
The prior art is difficult to effectively recognize video action under small sample conditions, especially in video scenarios where the environment and data are dynamically changed, and the model is difficult to maintain the ability to recognize old categories while quickly learning new categories.
The pre-trained contrast language-visual big model is used to extract video frame features through visual soft prompts and timing soft prompts, and combined with linear fusion and knowledge distillation technology, a small sample-class incremental video action recognition method is constructed, and video and text features are fused to alleviate the performance loss of the model in the old categories.
It significantly improves the recognition accuracy of new action categories, reduces the catastrophic forgetting of old categories by the model, and realizes efficient learning and recognition under small sample conditions.
Smart Images

Figure CN120561744A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and in particular relates to a method and device for small sample incremental video action recognition. Background Art
[0002] In recent years, deep neural networks have demonstrated remarkable performance in various computer vision and machine learning tasks, particularly when trained on large, labeled datasets. However, due to the long-tail distribution and Zipf's law, it is often difficult to obtain large amounts of supervised data upfront in real-world scenarios, and new categories may emerge dynamically over time. For example, in smart shelf monitoring systems, initially only common product categories (such as beverages, snacks, and daily necessities) can be recognized. Over time, merchants introduce new products (such as limited-edition beverages and seasonal items). For these new products, the system must quickly learn using only a small number of images while maintaining recognition accuracy for existing products. This requires the model to maintain its ability to recognize previously learned categories while continuously learning new categories using only a small number of labeled examples—a process known as small-sample incremental learning.
[0003] Incremental learning for small-shot classes typically consists of a base task and multiple sequential incremental tasks. In the base task, each class has a large number of labeled training examples for initial model construction, while in the incremental tasks, each class has only a few labeled examples for ongoing model training. When learning each task, the model can only be trained using the training data from the current task. After learning each task, the model needs to be tested on all previously seen classes. To address this issue, researchers have conducted in-depth research from multiple perspectives, such as pre-building a feature space that can adapt to incremental tasks, dynamically scaling the model architecture for incremental tasks, and stabilizing the feature topology space. In recent years, the Contrastive Language-Vision Grand Model, pre-trained on massive amounts of image-text data, has demonstrated outstanding zero-shot learning capabilities across a wide range of downstream tasks, including image segmentation and video understanding. Given its powerful knowledge transfer and generalization capabilities, researchers are exploring the application of the Contrastive Language-Vision Grand Model to address small-shot incremental learning, achieving impressive performance.
[0004] However, the aforementioned research primarily focuses on the image domain, with insufficient attention paid to the more challenging problem of video action recognition. Unlike static image tasks, video action recognition is more susceptible to the dynamic evolution of the environment and data, leading to the problem of small sample class augmentation. This problem is particularly prominent in fields such as security monitoring, healthcare, and human-computer interaction. For example, a popular video website can upload 500 hours of video data covering a wide range of categories every minute. Developing video action recognition models that can efficiently learn from this continuous data without requiring retraining is crucial. In the era of large models, solving the problem of small sample class augmentation video action recognition using large comparative language-vision models in a parameter-efficient manner is an urgent challenge. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and device for small sample incremental video action recognition in response to the deficiencies of the existing technology.
[0006] To achieve the above objectives, the present invention adopts the following technical solutions: First, a method for small sample incremental video action recognition is provided, comprising the following steps: Pre-train a large model for comparing language and image pre-training, vectorize each frame in the video into a video frame feature vector, input the pre-trained image encoder, average the output of all frames to obtain the prior video features; construct a visual soft prompt vector, splice it with the video frame feature vector (the visual soft prompt vector is used as a prefix prompt), input the pre-trained image encoder to obtain video frame features, construct a temporal soft prompt vector, and splice it with the video frame features of all frames in the video in chronological order (the temporal soft prompt vector is used as a prefix prompt) to obtain a video feature vector, input the randomly initialized temporal encoder to obtain video features; linearly fuse the video features and the prior video features; The pre-trained text encoder is used to obtain text prototypes for all categories (the text labels of the video categories are input into the text encoder to extract the text prototypes of the categories). The probability distribution of the similarity between the linearly fused video features and the text prototypes of each category and the cross entropy loss of the corresponding true labels of the videos are calculated, as well as the knowledge distillation loss between the linearly fused video features and the prior video features. The weighted addition is used as the loss function. The parameters of the image encoder and text encoder are frozen, and the visual soft prompt vector, temporal soft prompt vector, and temporal encoder are trained. When identifying the video to be tested, the probability distribution of the similarity between the linearly fused video features and the text prototypes of each category is calculated, and the category label corresponding to the maximum probability value is selected.
[0007] Furthermore, the temporal encoder is constructed based on a transformer structure.
[0008] Furthermore, the visual soft prompt vector and the temporal soft prompt vector are both composed of a plurality of soft prompt tokens, and the soft prompt tokens are randomly initialized.
[0009] Furthermore, the video features and the prior video features are linearly fused, and the weight of the video features is set to decrease as the number of performed learning tasks increases.
[0010] The video features after linear fusion are: ; in, is the video feature, is the prior video feature, For the current The fusion coefficient of the learning task, , represents the total number of learning tasks, 、 represents a hyperparameter.
[0011] Furthermore, the probability distribution of the similarity between the linearly fused video features and the text prototypes of each category is obtained by the following steps: calculating the cosine similarity between the linearly fused video features and the text prototypes of each category, and converting it into a probability distribution through a Softmax function.
[0012] Furthermore, the knowledge distillation loss of the linearly fused video features and the prior video features is obtained by the following steps: projecting the linearly fused video features into the prior video feature space, and performing residual connection to obtain the residually connected video features, and calculating the Euclidean distance between the prior video features and the residually connected video features as the knowledge distillation loss.
[0013] Furthermore, vectorizing each frame in the video into a video frame feature vector includes: dividing each frame in the video into multiple image blocks according to a fixed method, each image block is represented as a feature vector, and splicing the feature vectors in sequence to obtain the video frame feature vector.
[0014] In the second aspect, the present invention provides a small sample class incremental video action recognition device, comprising a memory and one or more processors, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute program data to implement the small sample class incremental video action recognition method described in the first aspect.
[0015] In a third aspect, the present invention provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the method for small sample incremental video action recognition described in the first aspect is implemented.
[0016] The present invention has the following beneficial effects: by employing visual and temporal soft cues, it can effectively extract spatiotemporal information from videos. This information is then further integrated with the prior features of the contrastive language-visual macromodel, thereby improving the generalization ability of the contrastive language-visual macromodel in downstream video action recognition tasks. Furthermore, by employing a progressive knowledge distillation method related to the learning task, it can effectively avoid performance loss on old action categories. The present invention's simple and flexible implementation significantly improves the recognition accuracy of new action categories while effectively mitigating the model's catastrophic forgetting of old categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Schematic diagram of the process of the small sample incremental video action recognition method of the present invention; Figure 2 This is a structural block diagram of the device for recognizing action from small sample size incremental videos according to the present invention; Figure 3 Schematic diagram of different learning tasks in the base class stage and incremental stage of the present invention; Figure 4 Schematic diagram of the process of sample input processing and progressive knowledge distillation of the present invention. DETAILED DESCRIPTION
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0019] It should be noted that, unless there is any conflict, the features in the following embodiments and implementations may be combined with each other.
[0020] like Figure 1 As shown, in an embodiment of the present invention, a small-sample incremental video action recognition method that integrates a pre-trained cross-modal large model and progressive knowledge distillation mainly includes training a pre-trained large model for contrast language images in the base class phase, and performing N learning tasks in the incremental phase to train visual soft prompts, temporal soft prompts, and a temporal encoder. The specific steps are as follows: (1) Construct a video feature extraction network to extract video features that integrate spatial and temporal information and prior information; wherein the network includes visual soft prompts, a temporal encoder, a temporal soft prompt, and a pre-trained image encoder. The step (1) includes the following sub-steps: (1.1) Use the open-source pre-trained Contrastive Language Image Pre-training Large Model (CLIP model) to initialize the image encoder and text encoder.
[0021] (1.2) Construct a temporal encoder based on the transformer structure and randomly initialize the temporal encoder.
[0022] (1.3) Constructing a visual soft cue vector, concatenating the visual soft cue vector with the input video frame to obtain a concatenated feature vector, and inputting the concatenated feature vector into an image encoder to obtain corresponding frame features. Step (1.3) includes: (1.3.1) First, construct the visual soft hint vector , and initialize randomly, where represents the set of real numbers, Indicates that the visual soft prompt vector has A soft prompt token, The length of the vector representing each soft hint token.
[0023] (1.3.2) Secondly, input video The i-th video frame is recorded as , where T indicates that the input video has a total of T frames, 、 Represents the length and width of the video frame, divides the video frame into M picture blocks with fixed size, and each picture block is represented by a length of The feature vector of the video frame is recorded as .
[0024] (1.3.3) Then, the constructed visual soft prompt vector and video frame feature vector Perform splicing to obtain the spliced feature vector .
[0025] (1.3.4) Finally, the concatenated feature vector is input into the image encoder In the video, the frame features corresponding to the video frame are obtained ,in Indicates the parameters of the image encoder.
[0026] (1.4) Constructing a temporal soft hint vector, concatenating the temporal soft hint vector with the video frame features obtained in step (1.3) to obtain a concatenated feature vector, and inputting the concatenated feature vector into a temporal encoder to obtain a video feature that integrates spatial and temporal information. Step (1.4) includes: (1.4.1) First, construct the timing soft hint vector , and initialize randomly, where Indicates that the timing soft vector has A soft prompt token, The length of the vector representing each soft hint token.
[0027] (1.4.2) Secondly, the timing soft vector All frame features of the input video To splice, Indicates that the input video has Frame, get the spliced feature vector (called video feature vector) .
[0028] (1.4.3) Finally, the concatenated feature vector is input into the temporal encoder In the video, the fusion time series information is obtained .
[0029] (1.5) Input the video frame into the image encoder constructed in (1.1) to obtain the corresponding frame features, and average all the frame features of the input video to obtain the corresponding prior video features.
[0030] First, obtain the features of each frame of the input video. Specifically, the i-th video frame of the input video is recorded as ,in 、 Represents the length and width of the video frame, divides the video frame into M picture blocks with fixed size, and each picture block is represented by a length of The feature vector of the video frame is recorded as , and the video frame Input image encoder In the video frame feature Secondly, all frame features of the video are averaged to obtain the prior video features .
[0031] (1.6) Fuse the video features obtained in (1.4) and (1.5) to obtain the final video features.
[0032] like Figure 3 As shown, the video features obtained by (1.4) And the prior video features obtained by (1.5) , perform linear fusion to obtain the final video features , which is: ; ; in, For the current The fusion coefficient of the learning task, Indicates that there are a total of A learning task, Represents a hyperparameter that controls the degree of coefficient attenuation. represents another hyperparameter.
[0033] (2) Construct a text feature extraction network to extract the text prototype of the video category.
[0034] (3) Calculate the cosine similarity and probability value between the video features obtained in step (1) and all category text prototypes obtained in step (2).
[0035] (4) Use the labeled training data of the current task for iterative training, such as Figure 4 As shown, during the training process, the parameters of the image encoder and the text encoder are fixed, and the cross entropy loss function is calculated based on the probability value obtained in step (3). At the same time, the progressive knowledge distillation loss is calculated. With the goal of minimizing the sum of the cross entropy loss function and the knowledge distillation, the parameters in the visual soft prompt, temporal soft prompt, and temporal encoder are adjusted to obtain the trained visual soft prompt, temporal soft prompt, and temporal encoder. The step (4) includes the following sub-steps: (4.1) First, calculate the video features and category text prototypes The similarity between them is calculated and the corresponding probability value is obtained. The calculation method is as follows: ; in Represents the text prototype corresponding to the c-th category; Indicates a total of categories; Represents the input video The corresponding predicted category label, Indicates the probability value of the predicted category label being category label c; represents the cosine similarity, Represents the exponential function.
[0036] (4.2) Then, calculate the cross entropy loss for each training video, and the calculation formula is as follows: ; in, Represents the input video The corresponding true label, Takes 0 or 1.
[0037] (4.3) Secondly, calculate video features With prior video features Specifically, the video features after linear fusion are Projection to prior video features In space, with Establish the residual connection g, get the video features after the residual connection, and then calculate Distillation loss of video features after residual concatenation.
[0038] The projection calculation method is as follows: ,in, and are the parameters of the two fully connected layers, represents the GELU activation function, Denotes the residual coefficient. The distillation loss is calculated as: .
[0039] (4.4) Finally, the cross entropy loss is combined with the knowledge distillation loss to obtain the final loss corresponding to the input video: ,in Indicates the balance coefficient.
[0040] (5) Given a video to be tested, use the trained visual soft prompts, temporal soft prompts and temporal encoders, as well as the image encoder and text encoder to repeat steps (1) to (3) to calculate the probability values between the video to be tested and all category labels, and select the category label corresponding to the maximum probability value as the final category label of the current video to be tested.
[0041] The present invention also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 A small-sample incremental video action recognition method is provided.
[0042] The present invention also provides Figure 2 The one shown corresponds to Figure 1 Schematic diagram of the small sample incremental video action recognition device of the method. Figure 2 As mentioned above, at the hardware level, the small sample incremental video action recognition device includes a processor, an internal bus, a network interface, a memory and a non-volatile memory, and may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Of course, in addition to software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0043] Improvements to a technology can be clearly distinguished as either hardware improvements (for example, improvements to circuit structures such as diodes, transistors, and switches) or software improvements (improvements to process flows). However, with technological advancements, many process flow improvements can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using a hardware module. For example, a programmable logic device (PLD) is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually crafting integrated circuit chips, this programming is mostly performed using "logic compiler" software. This is similar to the software compiler used during program development. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, with VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog being the most commonly used. Those skilled in the art will also understand that simply by programming a method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0044] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include the following microcontrollers: ARC625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the memory control logic. Those skilled in the art will also appreciate that, in addition to implementing the controller purely in computer-readable program code, the controller can also be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, an embedded microcontroller, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing the various functions included therein can also be considered as structures within the hardware component. Alternatively, the means for implementing the various functions can be considered both a software module implementing the method and a structure within the hardware component.
[0045] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0046] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0047] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0048] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0049] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0050] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0051] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0052] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0053] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0054] The present invention may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.
[0055] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A small sample incremental video action recognition method, characterized by: The following steps are involved: Pre-training comparison language image pre-training large model; each frame in the video is vectorized into a video frame feature vector, input into the pre-trained image encoder, and the output of all frames is averaged to obtain the prior video features; Construct a visual soft cue vector, concatenate it with the video frame feature vector, input it into a pre-trained image encoder to obtain video frame features, construct a temporal soft cue vector, concatenate it with the video frame features of all frames in the video in chronological order to obtain a video feature vector, input it into a randomly initialized temporal encoder to obtain video features; perform linear fusion on the video features and the prior video features; A pre-trained text encoder is used to obtain text prototypes for all categories. The probability distribution of the similarity between the linearly fused video features and the text prototypes of each category is calculated, along with the cross-entropy loss between the true labels of the corresponding videos. The knowledge distillation loss between the linearly fused video features and the prior video features is also calculated. The weighted sum of these two losses is used as the loss function. The parameters of the image encoder and text encoder are frozen, and the visual soft cue vector, temporal soft cue vector, and temporal encoder are trained. When identifying the video to be tested, the probability distribution of the similarity between the linearly fused video features and the text prototypes of each category is calculated, and the category label corresponding to the maximum probability value is selected.
2. The method according to claim 1, characterized in that The temporal encoder is constructed based on the transformer structure.
3. The method according to claim 1, characterized in that The visual soft prompt vector and the temporal soft prompt vector are both composed of multiple soft prompt tokens, and the soft prompt tokens are randomly initialized.
4. The method according to claim 1, wherein The method of vectorizing each frame in the video into a video frame feature vector includes: dividing each frame in the video into multiple picture blocks according to a fixed method, representing each picture block as a feature vector, and sequentially splicing the feature vectors to obtain the video frame feature vector.
5. The method according to claim 1, wherein The video features and the prior video features are linearly fused, and the video features after linear fusion are for: ; in, is the video feature, is the prior video feature, For the current The fusion coefficient of the learning task, , represents the total number of learning tasks, 、 is a hyperparameter.
6. The method according to claim 1, characterized in that The probability distribution of the similarity between the linearly fused video features and the text prototypes of each category is obtained by the following steps: calculating the cosine similarity between the linearly fused video features and the text prototypes of each category, and converting it into a probability distribution through a Softmax function.
7. The method according to claim 1, characterized in that The knowledge distillation loss of the linearly fused video features and the prior video features is obtained by the following steps: projecting the linearly fused video features into the prior video feature space and performing residual connection to obtain the residually connected video features, and calculating the Euclidean distance between the prior video features and the residually connected video features as the knowledge distillation loss.
8. A small sample incremental video action recognition device, comprising a memory and one or more processors, characterized in that: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the small sample class incremental video action recognition method according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that A program is stored thereon, and when the program is executed by a processor, the small sample incremental video action recognition method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Dynamic gesture recognition method, system and equipment and medium
CN116524593A
Video identification method and device, and storage medium
CN116740596A
Small sample class increment image classification method based on cue word fine tuning and feature playback
CN117746140A
Video abnormal event detection method based on prompt learning and multi-scale time sequence fusion
CN118918506A
Cross-language information sorting method and device based on semantic space clustering prompt, equipment and storage medium
CN119166780A