Video segmentation method and system, chip, display panel and device, and storage medium

By adopting a method based on historical frame attention matching in the video stream semantic segmentation algorithm, using downsampling and upsampling branch cross-processing and attention matching output, the problem of insufficient fusion performance of multi-history frame storage and multi-scale feature in the existing algorithm is solved, and a more efficient video stream semantic segmentation effect is achieved.

CN119964059APending Publication Date: 2025-05-09CHIPONE TECHNOLOGY (BEIJING) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510433969.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing video stream semantic segmentation algorithm faces the problem of low performance of multi-hierarchical frame storage and multi-scale feature fusion in frames, resulting in poor resource consumption and performance.

Method used

A video segmentation method based on historical frame attention matching is adopted, and multi-scale features are integrated through downsampling and upsampling branch cross-processing, and the semantic segmentation graph is output using attention matching.

Benefits of technology

The semantic segmentation effect of video streams is optimized, the multi-scale feature extraction performance and semantic segmentation efficiency are improved, and resource consumption is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964059A_ABST
    Figure CN119964059A_ABST
Patent Text Reader

Abstract

The invention discloses a video segmentation method and system, a chip, a display panel and device, and a storage medium. The video segmentation method provided by the embodiment of the invention comprises the following steps: acquiring a current video frame; performing down-sampling on the current video frame to obtain a feature under a first resolution; performing up-sampling on the current video frame to obtain features under a second resolution; fusing the feature under the second resolution into the feature under the first resolution to obtain a first feature corresponding to the down-sampling; fusing the feature under the first resolution into the feature under the second resolution to obtain a second feature corresponding to the up-sampling; and obtaining a semantic segmentation map according to the first feature and the second feature. According to the video segmentation method and system, the chip, the display panel and device and the storage medium in the embodiment of the invention, the semantic segmentation effect of the video stream is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of graphic display technology, and in particular to a video segmentation method and system, a chip, a display panel and device, and a storage medium. Background Art

[0002] Video stream semantic segmentation is a key technology in the field of computer vision, which involves classifying pixels in video frames to identify different objects and scenes. With the development of deep learning technology, video stream semantic segmentation has achieved significant improvements in accuracy and has been more widely used.

[0003] Existing video stream semantic segmentation algorithms often face problems such as storage of multiple historical frames and low performance of multi-scale feature fusion within frames. These problems seriously affect the algorithm resource consumption and performance.

[0004] Therefore, it is hoped that there will be a new video stream semantic segmentation solution that can overcome the above problems. Summary of the invention

[0005] In view of the above problems, the purpose of the present invention is to provide a video segmentation method and system, chip, display panel and device, storage medium, and especially a video segmentation method based on historical frame attention matching, so as to optimize the effect of semantic segmentation of video stream.

[0006] According to one aspect of the present invention, a video segmentation method is provided, comprising: obtaining a current video frame; downsampling the current video frame to obtain features at a first resolution; upsampling the current video frame to obtain features at a second resolution; integrating features at the second resolution into features at the first resolution to obtain first features corresponding to the downsampling; integrating features at the first resolution into features at the second resolution to obtain second features corresponding to the upsampling; and obtaining a semantic segmentation map based on the first features and the second features.

[0007] Optionally, the video segmentation method also includes: obtaining historical frames after semantic segmentation; performing feature extraction on the historical frames after semantic segmentation and the current video frame to obtain historical frame features and current video frame features respectively; fusing the historical frame features and the current video frame features to obtain fused features; and obtaining a semantic segmentation map corresponding to the current video frame based on the fused features.

[0008] Optionally, the video segmentation method further includes: encoding the historical frames after the semantic segmentation and the current video frame through a multi-scale residual module to generate corresponding multi-scale features.

[0009] Optionally, the video segmentation method further includes: performing feature mapping on the first feature and the second feature to obtain a feature vector of the current video frame; and obtaining the semantic segmentation map according to the feature vector of the current video frame.

[0010] Optionally, the video segmentation method further includes: obtaining historical frame features based on the historical frames after semantic segmentation; obtaining a key vector and a value vector respectively based on the historical frame features; and obtaining the semantic segmentation map based on the key vector and the value vector.

[0011] Optionally, the video segmentation method also includes: performing feature mapping on the first feature and the second feature to obtain a feature vector of the current video frame; performing attention matching output based on the feature vector of the current video frame, the key vector and the value vector, and outputting the semantic segmentation map.

[0012] Optionally, the current video frame is downsampled and feature extraction is performed through a convolution layer to obtain features at the first resolution; the current video frame is upsampled and feature extraction is performed through a convolution layer to obtain features at the second resolution.

[0013] According to another aspect of the present invention, a video segmentation system is provided, including: an acquisition unit, used to acquire a current video frame; a downsampling unit, used to downsample the current video frame to obtain features at a first resolution; an upsampling unit, used to upsample the current video frame to obtain features at a second resolution; a fusion unit, the fusion unit is used to integrate the features at the first resolution into the features at the second resolution to obtain the first features corresponding to the downsampling; the fusion unit is also used to integrate the features at the first resolution into the features at the second resolution to obtain the second features corresponding to the upsampling; and a segmentation unit, used to obtain a semantic segmentation map based on the first features and the second features.

[0014] According to yet another aspect of the present invention, there is provided a chip, comprising: the video segmentation system as described above.

[0015] According to another aspect of the present invention, a display panel is provided, comprising: the video segmentation system as described above, wherein the display panel comprises at least one selected from a cathode ray tube display panel, a digital light processing display panel, a liquid crystal display panel, a light emitting diode display panel, an organic light emitting diode display panel, a quantum dot display panel, a Mirco-LED display panel, a Mini-LED display panel, a field emission display panel, a plasma display panel, an electrophoretic display panel and an electrowetting display panel.

[0016] According to another aspect of the present invention, there is provided a display device, comprising: the video segmentation system as described above; and a display panel, wherein the display panel is connected to the video segmentation system to display the current video frame according to the semantic segmentation map.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, wherein the program implements the video segmentation method as described above when executed by a processor.

[0018] The video segmentation method and system, chip, display panel and device, and storage medium provided by the present invention perform cross-processing on downsampling branches and upsampling branches, so that each branch introduces relevant features of other branches, thereby optimizing the semantic segmentation effect of the video stream.

[0019] Furthermore, the historical frame features and the current video frame features are fused, thereby allowing more effective integration and utilization of multi-scale information in the current video frame and the historical frames, thereby improving the multi-scale feature extraction performance of the video stream.

[0020] Furthermore, attention matching output is performed based on the feature vector of the current video frame, the key vector and value vector of the historical frames, which improves the semantic segmentation effect and efficiency.

[0021] Furthermore, the current frame encoding module and the historical frame encoding module can be implemented on the standard transformer architecture (an architecture that uses a self-attention mechanism to process video spatiotemporal information, aiming to classify each frame in the video pixel by pixel while considering continuity in the time dimension), with a simple structure, simple method and a wide range of applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which: Figure 1 A method flow chart of a video segmentation method according to an embodiment of the present invention is shown; Figure 2 A system architecture diagram of a video segmentation system according to an embodiment of the present invention is shown; Figure 3 A schematic diagram of the structure of a multi-scale enhancement module according to an embodiment of the present invention is shown; Figure 4 A schematic diagram of the structure of a current frame encoding module according to an embodiment of the present invention is shown; Figure 5 A schematic diagram of the structure of a historical frame encoding module according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0023] Various embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. In each of the accompanying drawings, identical elements are represented by identical or similar reference numerals. For the sake of clarity, the various parts in the accompanying drawings are not drawn to scale. In addition, some well-known parts may not be shown in the drawings.

[0024] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. Many specific details of the present invention, such as the structure, materials, dimensions, processing technology and techniques of the components are described below to provide a clearer understanding of the present invention. However, as those skilled in the art will appreciate, the present invention may be implemented without following these specific details.

[0025] It should be understood that when describing the structure of a component, when a layer or a region is referred to as being "on" or "over" another layer or another region, it may mean that it is directly on the other layer or another region, or that other layers or regions are included between it and the other layer or another region. Moreover, if the component is turned over, the layer or a region will be "below" or "beneath" another layer or another region.

[0026] Figure 1 FIG. 4 is a flow chart of a method for video segmentation according to an embodiment of the present invention. Figure 1 As shown, the video segmentation method according to the embodiment of the present invention comprises the following steps: In step S101, the current video frame is obtained; Get the current video frame (current frame). In the video segmentation method, the current video frame is, for example, a video image (specific frame) corresponding to a certain moment that the algorithm is to process (analyze, segment, or mark).

[0027] In step S102, down-sampling the current video frame to obtain features at a first resolution; Downsample the current video frame to obtain features at the first resolution. Downsampling, for example, refers to a process of compressing the amount of data by reducing the resolution (spatial dimension) of the video frame to extract high-level (larger range) semantic features. Optionally, downsample the current video frame and perform feature extraction through a convolutional layer to obtain features at the first resolution.

[0028] In step S103, upsampling the current video frame to obtain features at a second resolution; The current video frame is upsampled to obtain features at the second resolution. Upsampling refers to restoring the resolution to generate a fine segmentation result. Optionally, the current video frame is upsampled and feature extraction is performed through a convolutional layer to obtain features at the second resolution.

[0029] In step S104, the features at the second resolution are integrated into the features at the first resolution to obtain the first features corresponding to the downsampling; The features at the second resolution are integrated into the features at the first resolution to obtain the first features corresponding to the downsampling, that is, the relevant features of another scale (in the upsampling) are introduced into the downsampling branch.

[0030] In step S105, the features at the first resolution are integrated into the features at the second resolution to obtain second features corresponding to the upsampling; The features at the first resolution are integrated into the features at the second resolution to obtain the second features corresponding to the upsampling, that is, the relevant features of another scale (in the downsampling) are introduced into the upsampling branch.

[0031] In step S106, a semantic segmentation map is obtained according to the first feature and the second feature.

[0032] The semantic segmentation map of the current video frame is obtained according to the first feature and the second feature. The obtained semantic segmentation map, for example, assigns a semantic label to each pixel in the current video frame, indicating the category to which the pixel belongs. The semantic segmentation map has the same resolution as the original image, and the value of each pixel is no longer color information, but a corresponding category label.

[0033] In an optional embodiment of the present invention, the video segmentation method further includes the following steps: Obtain the historical frame after semantic segmentation; optionally, the current video frame that has completed semantic segmentation is used as the historical frame of the subsequent video frame.

[0034] Feature extraction is performed on the historical frames and the current video frame after semantic segmentation to obtain historical frame features and current video frame features respectively; the current video frame features include, for example, a first feature and a second feature.

[0035] Fuse the historical frame features with the current video frame features to obtain fused features; The semantic segmentation map corresponding to the current video frame is obtained based on the fusion features.

[0036] In the above embodiment of the present invention, features of different resolutions are captured by downsampling and upsampling branches, and feature extraction is performed through the convolution layer; then the upper and lower branch scales (downsampling and upsampling) are cross-processed to ensure that each branch introduces relevant features of another scale, thereby optimizing the video stream semantic segmentation effect. Furthermore, the historical frame features and the current video frame features are fused, allowing multi-scale information to be more effectively integrated and utilized in the current video frame and the historical frame, thereby improving the video stream multi-scale feature extraction performance.

[0037] Optionally, the video segmentation method further includes: encoding the historical frames and the current video frames after semantic segmentation through a multi-scale residual module to generate corresponding multi-scale features.

[0038] In an optional embodiment of the present invention, the video segmentation method further includes: performing feature mapping on the first feature and the second feature to obtain a feature vector of the current video frame; and obtaining a semantic segmentation map based on the feature vector of the current video frame. Optionally, obtaining historical frame features based on the historical frames after semantic segmentation; obtaining a key vector and a value vector respectively based on the historical frame features; and obtaining a semantic segmentation map based on the key vector and the value vector. Optionally, performing feature mapping on the first feature and the second feature to obtain a feature vector of the current video frame; performing attention matching output based on the feature vector, key vector and value vector of the current video frame, and outputting a semantic segmentation map.

[0039] According to another aspect of the present invention, a video segmentation system is provided, which is used to implement the above-mentioned video segmentation method. The video segmentation system includes an acquisition unit, a down-sampling unit, an up-sampling unit, a fusion unit and a segmentation unit.

[0040] Specifically, the acquisition unit is used to acquire the current video frame.

[0041] The down-sampling unit is used to down-sample the current video frame to obtain features at a first resolution.

[0042] The up-sampling unit is used to up-sample the current video frame to obtain features at a second resolution.

[0043] The fusion unit is used to integrate the features at the first resolution into the features at the second resolution to obtain the first features corresponding to the downsampling. The fusion unit is also used to integrate the features at the first resolution into the features at the second resolution to obtain the second features corresponding to the upsampling.

[0044] The segmentation unit is used to obtain a semantic segmentation map according to the first feature and the second feature.

[0045] The video segmentation method and video segmentation system of the present application are described in detail below in conjunction with specific embodiments.

[0046] Figure 2 FIG. 4 shows a system architecture diagram of a video segmentation system according to an embodiment of the present invention. Figure 2As shown, the system framework is, for example, a historical frame attention matching (Historical Frame Attention Match, HFAM) algorithm framework, which includes a multi-scale residual module / residual encoder (Res Encoder), a multi-scale enhancement module (MSE), a historical frame encoding module (His-Frame Encoder), a current frame (current video frame) encoding module (Cur-Frame Encoder), a storage matching attention / memory attention module (MA Module) and a residual decoder (Res Decoder).

[0047] like Figure 2 As shown in the figure, the algorithm simultaneously inputs the historical frames after semantic segmentation and the current (video) frame to be segmented. After the historical frames are processed by the residual encoder, multi-scale enhancement module and historical frame encoding module in sequence, the processing results (features, etc.) are sent to the memory attention module. After the current video frame is processed by the residual encoder, multi-scale enhancement module and current frame encoding module in sequence, the processing results (features, etc.) are sent to the memory attention module. The memory attention module processes the processing results of the received historical frames and the processing results of the current video frame, and the processed results are passed through the residual decoder to output the (semantic) segmentation map.

[0048] Figure 3 FIG. 4 shows a schematic diagram of the structure of a multi-scale enhancement module according to an embodiment of the present invention. Figure 3 As shown, the multi-scale enhancement module (multi-scale feature fusion module) includes two branches, one for downsampling and the other for upsampling.

[0049] For the downsampling branch, the convolution layer is used to extract features of the current video frame. The multi-scale enhancement module also performs function operations (such as Sigmoid function (S-type function)) and layer normalization on the downsampling branch.

[0050] For the upsampling branch, the convolution layer is used to extract features of the current video frame. The multi-scale enhancement module also performs function operations (such as Sigmoid function (S-type function)) and layer normalization on the downsampling branch.

[0051] The multi-scale enhancement module also cross-processes the down-sampling branch and the up-sampling branch to ensure that each branch introduces relevant features of another scale (branch). In the figure, C (Cat) refers to the splicing of features along the channel dimension.

[0052] Optionally, the multi-scale enhancement module has at least one downsampling branch and at least one upsampling branch. The multi-scale enhancement module amplifies the low-resolution features output by the downsampling branch to the same or similar resolution as the high-resolution features output by the upsampling branch through an upsampling operation, and reduces the high-resolution features output by the upsampling branch to the same or similar resolution as the low-resolution features output by the downsampling branch through a downsampling operation, and fuses the amplified low-resolution features with the original high-resolution features through a convolution layer, and fuses the reduced high-resolution features with the original low-resolution features through a convolution layer. The fused features are reprocessed through the respective up and down sampling branches to ensure full interaction and integration of multi-scale information.

[0053] Figure 4 FIG. 4 shows a schematic diagram of the structure of a current frame encoding module according to an embodiment of the present invention. Figure 4 As shown in the figure, the current frame encoding module adopts a standard transformer architecture (a deep learning architecture based on the self-attention mechanism), embeds a residual module for feature mapping, and maps it into a Q (Query, query vector) value in the second layer of the residual module.

[0054] Specifically, the current frame encoding module according to an embodiment of the present invention includes image block embedding (PatchEmbedding), multi-head self-attention (Multi-Head Self-Attention, MSA), multi-criteria analysis (Multi-Criteria Analysis, MCA), residual block (ResBlock) and feedforward neural network (FFN).

[0055] Combination Figure 2 and Figure 4 As shown in the figure, the current frame encoding module receives the processing result of the multi-scale enhancement module on the current video frame. The processing result is processed by image block embedding, multi-head self-attention, multi-criteria analysis, residual block, feedforward neural network and residual block to obtain the Q value.

[0056] Figure 5 FIG. 4 shows a schematic diagram of the structure of a historical frame encoding module according to an embodiment of the present invention. Figure 5 As shown in Figure 1, the historical frame encoding module adopts the standard transformer architecture, but adds two layers of residual blocks (ResBlock) after the feedforward neural network (FFN) to obtain K (Key, key vector) and V (Value, value vector) respectively.

[0057] Specifically, the historical frame encoding module according to an embodiment of the present invention includes image block embedding (PatchEmbedding), multi-head self-attention (Multi-Head Self-Attention, MSA), multi-criteria analysis (Multi-Criteria Analysis, MCA), residual block (ResBlock) and feedforward neural network (FFN).

[0058] Combination Figure 2 and Figure 5 As shown in the figure, the historical frame encoding module receives the processing results of the historical frame by the multi-scale enhancement module. The processing results are processed by image block embedding, multi-head self-attention, multi-criteria analysis, residual block and feedforward neural network, and then processed by residual block to obtain K value and V value respectively.

[0059] In an optional embodiment of the present invention, the Q value obtained by the current frame encoding module and the K value and V value obtained by the historical frame encoding module are input to the memory attention module (MA Module) for attention matching and output. Optionally, the specific formula is as follows: F = MSA (Q, K, V) Among them, F is the feature value after query, MSA (Mult-head Self-Attention) is multi-head self-attention, and Q is the query vector. In the attention mechanism, Q represents the feature vector of the current video frame, which is used to "ask" which historical or contextual information needs to be paid attention to; K is the key vector and V is the value vector. K and V usually come from historical frames or other areas of the current video frame. Q calculates the similarity weight with K and weighted aggregates the information in V.

[0060] The memory attention module calculates the similarity between the query vector (Q) and the key vector (K) to obtain the attention weight, and performs weighted summation on the value vector (V) based on the attention weight, and then outputs the final semantic segmentation result (semantic segmentation map).

[0061] In a specific embodiment, taking the current video frame display screen as a person running on the grass as an example, the historical frame display screen is the person running on the grass at the previous moment. According to the processing results of the historical frame after being processed by the multi-scale enhancement module, the historical frame encoding module, the memory attention module, etc. and the processing results of the current video frame after being processed by the multi-scale enhancement module, the current frame encoding module, the memory attention module, etc., a semantic segmentation map corresponding to the current video frame is obtained. In the semantic segmentation map of the current video frame, the person, the person's shadow, and the grass (background) in the current video frame display screen can be well segmented.

[0062] According to another aspect of the present invention, a (display) chip is provided, which includes the video segmentation system as described above.

[0063] According to another aspect of the present invention, a display panel is provided. The display panel includes the video segmentation system as described above. The display panel includes at least one selected from a cathode ray tube display panel, a digital light processing display panel, a liquid crystal display panel, a light emitting diode display panel, an organic light emitting diode display panel, a quantum dot display panel, a Mirco-LED display panel, a Mini-LED display panel, a field emission display panel, a plasma display panel, an electrophoretic display panel, or an electrowetting display panel.

[0064] According to another aspect of the present invention, a display device is provided. The display device comprises the video segmentation system as described above and a display panel. The display panel is connected to the video segmentation system to display a current video frame according to a semantic segmentation map.

[0065] According to another aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, wherein the program implements the video segmentation method as described above when executed by a processor. The computer storage medium of the embodiment of the present disclosure may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. The computer program code for performing the operation of the embodiment of the present disclosure may be written in one or more programming languages ​​or a combination thereof, and the programming language includes an object-oriented programming language (such as Java, Smalltalk, C++), and also includes a conventional procedural programming language such as "C" language or a similar programming language.

[0066] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0067] According to the embodiments of the present invention as described above, these embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made based on the above description. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can make good use of the present invention and the modified use based on the present invention. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A video segmentation method, comprising: Get the current video frame; Downsampling the current video frame to obtain features at a first resolution; Upsampling the current video frame to obtain features at a second resolution; Integrate the features at the second resolution into the features at the first resolution to obtain the first features corresponding to the downsampling; Integrating the features at the first resolution into the features at the second resolution to obtain second features corresponding to the upsampling; and A semantic segmentation map is obtained according to the first feature and the second feature.

2. The video segmentation method according to claim 1, wherein: The video segmentation method also includes: Get the historical frames after semantic segmentation; Performing feature extraction on the semantically segmented historical frames and the current video frame to obtain historical frame features and current video frame features respectively; Fusing the historical frame features and the current video frame features to obtain fused features; and A semantic segmentation map corresponding to the current video frame is obtained according to the fusion features.

3. The video segmentation method according to claim 2, wherein: The video segmentation method also includes: The historical frames after the semantic segmentation and the current video frame are encoded through a multi-scale residual module to generate corresponding multi-scale features.

4. The video segmentation method according to claim 1, wherein: The video segmentation method also includes: Performing feature mapping on the first feature and the second feature to obtain a feature vector of the current video frame; and The semantic segmentation map is obtained according to the feature vector of the current video frame.

5. The video segmentation method according to claim 1, wherein: The video segmentation method also includes: Obtain historical frame features based on the historical frames after semantic segmentation; Obtaining a key vector and a value vector respectively according to the historical frame features; and The semantic segmentation map is obtained according to the key vector and the value vector.

6. The video segmentation method according to claim 5, wherein: The video segmentation method also includes: Performing feature mapping on the first feature and the second feature to obtain a feature vector of the current video frame; Attention matching is performed according to the feature vector of the current video frame, the key vector and the value vector, and the semantic segmentation map is output.

7. The video segmentation method according to claim 1, wherein: Down-sampling the current video frame and performing feature extraction through a convolutional layer to obtain features at the first resolution; The current video frame is upsampled and feature extracted through a convolutional layer to obtain features at the second resolution.

8. A video segmentation system, comprising: An acquisition unit, used for acquiring a current video frame; A down-sampling unit, configured to down-sample the current video frame to obtain features at a first resolution; an upsampling unit, configured to upsample the current video frame to obtain features at a second resolution; a fusion unit, the fusion unit being used to merge the features at the second resolution into the features at the first resolution to obtain the first features corresponding to the downsampling; the fusion unit is further used to merge the features at the first resolution into the features at the second resolution to obtain the second features corresponding to the upsampling; and A segmentation unit is used to obtain a semantic segmentation map according to the first feature and the second feature.

9. A chip, comprising: The video segmentation system as claimed in claim 8.

10. A display panel, comprising: The video segmentation system as claimed in claim 8, Among them, the display panel includes at least one selected from cathode ray tube display panel, digital light processing display panel, liquid crystal display panel, light emitting diode display panel, organic light emitting diode display panel, quantum dot display panel, Mirco-LED display panel, Mini-LED display panel, field emission display panel, plasma display panel, electrophoretic display panel and electrowetting display panel.

11. A display device, comprising: The video segmentation system as claimed in claim 8; as well as A display panel is connected to the video segmentation system to display the current video frame according to the semantic segmentation map.

12. A computer-readable storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the video segmentation method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and device for segmenting video object and network model training method

    CN113506316A

  • Video instance segmentation method and segmentation device based on space-time memory information

    CN114241388A

  • Multi-scale fusion remote sensing image semantic segmentation method and system

    CN115512103A

  • Real-time semantic segmentation method and system for dual-resolution interactive attention

    CN118038053A

  • Lightweight multi-target tracking method based on attention mechanism

    CN118172387A