Chinese opera stage scene segmentation method based on lightweight backbone structure
Through wavelet transformation, fractal dimensions and information entropy optimization computing resources, combined with the feature propagation method of nonlinear dynamics model, the problems of resource waste and information loss in image segmentation of complex stage scenes are solved, achieving a more efficient and accurate segmentation effect.
Patent Information
- Application Number
- CN202510554451.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art cannot maintain both global and local information in complex stage scene image segmentation, the allocation of computing resources is inefficient and resource waste, and the feature transfer and fusion methods are limited, resulting in poor segmentation effect.
Using a method based on lightweight backbone structure, multi-scale features are extracted through wavelet transformation, fractal dimensions and information entropy are calculated to adjust computing resources, and feature information propagation is optimized by combining nonlinear dynamic models, and multi-level feature fusion is carried out.
It improves the accuracy and efficiency of image segmentation, can better capture details in complex stage scenes, avoid information loss, and improves the boundary accuracy and area recognition accuracy of segmentation results.
Smart Images

Figure CN120472160A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a method for segmenting opera stage scenes based on a lightweight backbone structure. Background Art
[0002] Currently, image segmentation methods for complex stage scenes usually use traditional image processing techniques such as threshold segmentation and edge detection. Although these methods can work effectively in simple scenes, their performance is greatly reduced when processing images with complex backgrounds and details such as opera stages. In particular, traditional methods are often unable to maintain global and local information at the same time, and are prone to ignoring details or the blurred boundaries between background and foreground. The shortcoming of this technical solution is the lack of in-depth exploration of the multi-scale features of the image, which makes the segmentation results in complex scenes often not accurate enough.
[0003] In addition, the resource allocation method in the existing technology is usually static, and fixed computing resources are allocated throughout the image. The amount of resources allocated is basically the same regardless of the background area of the image or the foreground area with rich details. However, in reality, the demand for computing resources in different areas of the image varies greatly. Complex background and detailed areas require more computing resources, while simple background areas can reduce resource consumption. The existing static resource allocation method not only leads to low computational efficiency, but also leads to resource waste, which makes it difficult to achieve the best effect of segmentation tasks in practical applications.
[0004] Furthermore, existing feature transfer and fusion methods have limited effects when processing complex scenes; most traditional methods transfer features linearly from low layers to high layers, but this method is prone to information loss or blurred details when faced with highly nonlinear scenes; for example, edge detection or threshold methods often lose fine textures or tiny object boundaries and cannot guarantee the accuracy of segmentation; existing technologies fail to combine nonlinear dynamic models well to optimize information flow, resulting in often unsatisfactory segmentation effects in scenes with dynamic changes or complex structures; therefore, the present invention proposes a method for segmenting opera stage scenes based on a lightweight backbone structure to address the shortcomings of the existing technology. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides a method for segmenting opera stage scenes based on a lightweight backbone structure, which solves the problems of insufficient feature extraction, inefficient allocation of computing resources, limited information transmission and insufficient segmentation accuracy in traditional methods under complex backgrounds.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for segmenting opera stage scenes based on a lightweight backbone structure, comprising the following steps:
[0007] S1, obtaining an opera stage scene image to be segmented and performing image preprocessing;
[0008] S2. Decomposing the image by wavelet transform to obtain high-frequency components and low-frequency components of the scale;
[0009] S3, calculating the fractal dimensions of the high-frequency components and the low-frequency components of each scale of the image, and adjusting subsequent calculation parameters based on the fractal dimensions;
[0010] S4. Determine the size and distribution of the convolution kernel according to the fractal dimension, and perform feature extraction;
[0011] S5. Calculate the distribution information of the image and adjust the computing resources of each region in the feature extraction process based on the information entropy;
[0012] S6, optimize the extracted features by combining nonlinear dynamics model and use nonlinear activation function to adjust information propagation;
[0013] S7. Fuse the optimized features and output the segmentation results of the opera stage scene.
[0014] The present invention also provides a drama stage scene segmentation system based on a lightweight backbone structure, comprising:
[0015] An input module is used to receive the opera stage scene image to be segmented and perform preprocessing;
[0016] A feature extraction module is used to extract multi-scale features of the input image using wavelet transform and calculate the fractal dimension;
[0017] Computing resource allocation module, used to adjust the computing resources of each area according to the calculation results of fractal dimension and information entropy;
[0018] Information propagation optimization module, which is used to optimize information propagation by combining nonlinear dynamics and adjust feature information transmission through nonlinear activation functions;
[0019] The fusion and output module is used to fuse the optimized feature information and output the segmentation results of the opera stage scene.
[0020] The present invention provides a method for segmenting opera stage scenes based on a lightweight backbone structure.
[0021] Beneficial effects:
[0022] 1. The present invention adopts a technical solution that combines multi-scale feature extraction and wavelet transform to achieve the effect of comprehensively extracting image detail features; compared with the technical solution in the prior art that only relies on single-scale feature extraction, the present invention can capture the multi-level complex details in the opera stage scene, especially when the texture and background change greatly, solving the deficiency of traditional methods that cannot simultaneously maintain global and local information.
[0023] 2. The present invention achieves the effect of dynamically allocating computing resources according to image content by introducing computing resource allocation technology based on fractal dimension and information entropy. Compared with the fixed resource allocation scheme in the prior art, this scheme can automatically adjust resources according to the information complexity of different areas, solving the problems of resource waste and low processing efficiency, and ensuring that high-information areas are processed first.
[0024] 3. The present invention adopts a nonlinear dynamic model to optimize information propagation, and combines a nonlinear activation function to adjust the feature information transmission path, thereby achieving the effect of enhancing the transmission of key information; compared with traditional linear information propagation technology, the present invention can effectively improve the perception of details in complex scenes, avoid the problems of information loss or over-smoothing, and thus make the segmentation results more accurate and detailed.
[0025] 4. The present invention achieves a refined segmentation effect by combining multi-level feature fusion with dynamic threshold adjustment. Compared with the simple global threshold method in the prior art, the present invention can adaptively adjust the segmentation threshold according to local contrast, avoiding segmentation errors caused by illumination and background interference, thereby greatly improving the boundary precision and region recognition accuracy of the segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a flow chart of the method of the present invention;
[0027] Figure 2 This is a system architecture diagram of the present invention. DETAILED DESCRIPTION
[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0029] See also Figure 1 The embodiment of the present invention provides a method for segmenting opera stage scenes based on a lightweight backbone structure, comprising the following steps:
[0030] S1, obtaining an opera stage scene image to be segmented and performing image preprocessing;
[0031] In the present invention, step S1 mainly involves acquiring and preprocessing the opera stage scene image to be segmented; this process, as the first step of the entire segmentation method, plays a vital role in subsequent image analysis and feature extraction; image preprocessing is to ensure that subsequent processing can proceed smoothly, while improving the accuracy and efficiency of image segmentation.
[0032] First, the images to be segmented enter the process through the image input module; these images may come from different sources, such as stage photos, video frames, or other digital image data; due to differences in image quality, they often contain problems such as noise and uneven brightness, so image preprocessing is a necessary step.
[0033] In one possible implementation, image preprocessing includes operations such as image denoising, normalization, and color space conversion. Image denoising aims to remove high-frequency noise from an image in order to extract effective scene features. Denoising methods may include Gaussian filtering, median filtering, and other techniques. Specifically, Gaussian filtering is a commonly used method for smoothing images. It blurs the image using a given Gaussian kernel function to reduce the impact of noise on the image. In addition, median filtering is a nonlinear filtering method that can effectively remove salt and pepper noise from an image, especially when processing edge and detail information.
[0034] In general, during preprocessing, images need to be normalized to map the value of each pixel to a standard range, usually [0, 1]. The purpose of this normalization is to eliminate brightness differences in the image due to different shooting equipment or environmental conditions, ensuring the stability of subsequent algorithms. Specifically, for a given image I, its normalized pixel value can be calculated using the following formula:
[0035]
[0036] Among them, I i,j Represents the pixel value of the i-th row and j-th column of the original image, min(I) and max(I) are the minimum and maximum pixel values in image l, respectively, I′ i,j is the normalized pixel value.
[0037] In another possible implementation, an important step in preprocessing is color space conversion. Since images in the RGB color space contain too much redundant information, especially the correlation between different color channels, conversion to other color spaces (such as YCbCr) can better perform image segmentation. The YCbCr color space helps to remove the impact of lighting changes on image processing by separating brightness information and chromaticity information.
[0038] Specifically, the formula for converting RGB color space to YCbCr color space is as follows:
[0039] Y=0.299R+0.587G+0.114B;
[0040] Cb=-0.168736R-0.331264G+0.5B;
[0041] Cr=0.5R-0.418688G-0.081312B;
[0042] Among them, R, G, and B are the pixel values of the red, green, and blue channels of the original image, respectively, Y represents brightness information, and Cb and Cr represent chrominance information; the converted YCbCr image can be used for subsequent feature extraction and analysis more clearly.
[0043] As an option, image preprocessing can also include edge enhancement processing of the image to improve the performance of edge features in subsequent processing; for example, the Sobel operator, Laplacian operator, etc. can be used for edge detection to better capture the contour information of the image in the subsequent convolution process.
[0044] In this embodiment, the goal of image preprocessing is to remove noise and balance the brightness and color of the image through technical means such as denoising, normalization, and color space conversion, so that the image is more suitable for subsequent processing steps, especially operations such as wavelet transform and feature extraction; this process can effectively improve image quality and lay a solid foundation for subsequent segmentation tasks.
[0045] S2. Decomposing the image by wavelet transform to obtain high-frequency components and low-frequency components of the scale;
[0046] In step S2 of the present invention, wavelet transform is mainly used to perform multi-scale decomposition on the image to obtain different frequency components of the image to extract features and provide a basis for subsequent segmentation; since the opera stage scene contains rich colors, textures and complex background information, wavelet transform can effectively separate important low-frequency structural information and high-frequency detail information, making subsequent feature extraction and region segmentation more accurate.
[0047] In step S1, the image has been preprocessed, including denoising, normalization, color space conversion and other operations, providing a standardized input image for wavelet transform; based on this, the goal of step S2 is to perform wavelet transform on the input image and decompose it into frequency components of different scales, so as to better analyze the image content.
[0048] In this embodiment, the wavelet transform adopts a two-dimensional discrete wavelet transform (2D-DWT). Its basic process includes filtering in the horizontal, vertical, and diagonal directions, combined with a downsampling operation to reduce data redundancy. Mathematically, the core formula of the wavelet transform can be expressed as follows:
[0049]
[0050] Among them, W ψ (a, b) are wavelet transform coefficients, indicating the signal f(t) in the wavelet basis ψ a,b (t); f(t) is the original input signal; a is the scale parameter, which controls the time scalability of the wavelet; b is the translation parameter, which controls the position of the wavelet on the time axis; t is the time variable; ψ a,b (t) is the scaled and translated wavelet basis function, which is defined as follows:
[0051]
[0052] In general, the wavelet transform is implemented using a decomposition filter bank, including a low-pass filter G and a high-pass filter H for decomposition. The specific calculation is as follows:
[0053] LL(i,j)=∑ m ∑ n f(m,n)G(i-2m)G(j-2n);
[0054] LH(i,j)=∑ m ∑ n f(m,n)G(i-2m)H(j-2n);
[0055] HL(i,j)=∑ m ∑ n f(m,n)H(i-2m)G(j-2n);
[0056] HH(i,j)=∑ m ∑ n f(m,n)H(i-2m)H(j-2n);
[0057] Among them, LL(i,j) is the low-frequency (approximation) component, which contains the large-scale information of the image; LH(i,j) is the horizontal high-frequency component (horizontal details), which mainly reflects the vertical edge information; HL(i,j) is the vertical high-frequency component (vertical details), which mainly reflects the horizontal edge information; HH(i,j) is the diagonal high-frequency component (diagonal details), which contains complex high-frequency details; f(m,n) is the pixel value of the original input image; G is a low-pass filter (low-pass filter) used to extract low-frequency information; H is a high-pass filter (high-pass filter) used to extract high-frequency information; i,j are pixel coordinates; m,n are index variables of the convolution operation.
[0058] In one possible implementation, the wavelet basis function may be Daubechies wavelet (dbN), where N represents the order of the wavelet. Daubechies wavelet is a compactly supported orthogonal wavelet suitable for signal and image processing. Its mathematical expression is as follows:
[0059] ψ dbN (t)=∑ k h k ψ(2t-k);
[0060] Among them, ψ dbN (t) is the Daubechies wavelet basis function; h k is the Daubechies filter coefficient; k is the filter index; ψ(2t-k) is the scaled and translated wavelet mother function; N is the order of the Daubechies wavelet (such as db2, db4, db8, etc.).
[0061] As an alternative, another wavelet basis can use the Hart wavelet (Haar), which is defined as follows:
[0062]
[0063] Among them, ψ Haar (t) is the Haar wavelet basis function;
[0064] This function is a simple square wave function:
[0065] In the interval The value inside is 1; in the interval The value is -1 if it is internal; otherwise, the value is 0;
[0066] Haar wavelet is the simplest wavelet with fast calculation speed and is suitable for fast decomposition and feature extraction.
[0067] Specifically, wavelet transform can apply multi-level decomposition to perform a deeper multi-scale analysis of the image; for example, in the case of two-level wavelet decomposition, the original image is first decomposed to obtain four sub-bands LL1, LH1, HL1, and HH1; then LL1 is subjected to wavelet transform again to further extract the LL2, LH2, HL2, and HH2 sub-bands.
[0068] In some embodiments, the level of wavelet decomposition may be adjusted to adapt to scenes of different complexities; for example, in a stage scene, the high-frequency features of lighting and costumes are more important, so a level 3 or 4 wavelet transform may be selected to more accurately extract these detailed information; and for situations where the background is simpler, a level 1 or 2 wavelet transform may be used to reduce computational complexity.
[0069] In addition, in order to improve computational efficiency, the fast wavelet transform (FWT) method can be used in some embodiments to reduce the amount of computation through recursive calculation; FWT adopts a convolution-downsampling method and uses a discrete filter bank for transformation, so that the computational complexity is reduced to O(N).
[0070] In one possible implementation, in order to avoid boundary effects, when performing wavelet transform, the boundary extension strategy may adopt a symmetric extension or a periodic extension method; for example, the symmetric extension method is defined as follows:
[0071] f(x)=f(-x), x<0;
[0072] Among them, this formula is used to deal with signal boundary problems and prevent boundary artifacts; f(x) is the input signal;
[0073] This method expands the signal in a mirror-symmetrical manner so that data at the boundary can be processed correctly, especially avoiding discontinuities during wavelet transformation.
[0074] In this embodiment, the low-frequency components after wavelet transformation can be used for global feature extraction, while the high-frequency components are mainly used for local detail analysis. According to actual application requirements, appropriate wavelet basis functions, decomposition levels and boundary processing strategies can be selected to optimize the feature extraction effect.
[0075] S3, calculating the fractal dimensions of the high-frequency components and the low-frequency components of each scale of the image, and adjusting subsequent calculation parameters based on the fractal dimensions;
[0076] In the aforementioned step S2, the original image has been decomposed into high-frequency components and low-frequency components of different scales through wavelet transform; the low-frequency components mainly contain the global structural information of the image, while the high-frequency components mainly reflect local texture and edge features; in this step S3, the fractal dimension of these components is further calculated to quantify the complexity characteristics of the image at different scales; this calculation is of great significance for subsequent feature extraction, image analysis and classification tasks.
[0077] Generally speaking, fractal dimension is an indicator used to measure the complexity of signals or images, which can reflect the self-similarity and degree of detail variation of the image. The larger the fractal dimension, the richer the detail information contained in the image, while the smaller the fractal dimension, the smoother the image and the simpler the structure. Therefore, in the image processing process, fractal dimension can be used as one of the key features to distinguish different categories of images.
[0078] In one possible implementation, this embodiment uses the box counting method to calculate the fractal dimension. This method divides the image into grids at different scales and counts the number of grids containing image features (such as edge pixels or textures) to estimate the fractal dimension. The mathematical expression is as follows:
[0079]
[0080] Where: D f is the fractal dimension; ∈ is the size of the grid, that is, the size of the divided box; N(∈) is the number of grids containing image features (such as edge pixels, non-zero pixels after binarization) when the grid size is ∈.
[0081] The specific calculation steps are as follows:
[0082] Step 1: Binarize the high-frequency components of different scales and convert them into binary images to extract edge and texture information;
[0083] Step 2: Use a grid of size ∈ to cover the entire image and count the number of grids N(∈) that contain at least one non-primitive pixel;
[0084] Step 3: Change the value of ∈ and calculate the corresponding logN(∈) and log(1 / ∈);
[0085] Step 4: Perform linear regression on the obtained data and calculate the slope D of the fitting line f , that is, the fractal dimension.
[0086] As an option, in some embodiments, the power spectrum method (Power Spectrum Method) can be used to calculate the fractal dimension; this method is based on the Fourier transform of the image and estimates the fractal dimension by calculating the power spectral density (PSD) of the image. Its mathematical formula is as follows:
[0087] P(f)~f -β ;
[0088] Where: P(f) is the power spectral density; f is the spatial frequency; β is the spectral index, and its value is related to the fractal dimension D f The relationship between them is:
[0089]
[0090] In this method, the image is first Fourier transformed to calculate its spectrum, and then the power spectrum density curve is fitted in the logarithmic coordinate to solve the spectrum index β, and finally the fractal dimension D is obtained. f .
[0091] In some embodiments, in order to improve the accuracy of the calculation, a multi-scale fractal analysis method can be combined; multi-scale fractal analysis can calculate the local fractal dimension at different scales, thereby providing a more detailed description of complexity; its mathematical expression is as follows:
[0092]
[0093] Where: D q is the generalized fractal dimension; q is the moment order, which is used to adjust the sensitivity of the measure to large or small structures; p i is the normalized measure value of the i-th grid, indicating the relative importance of the grid; ∈ is the grid size.
[0094] The main steps of this method include:
[0095] Calculate the grid distribution at different scales ∈;
[0096] Statistical probability measure p of each grid i ;
[0097] Select different moment orders q (usually including negative numbers, decimals and positive integers) for calculation;
[0098] Determining the generalized fractal dimension D by curve fitting q .
[0099] In some embodiments, the fractal dimensions may be calculated for the high-frequency component and the low-frequency component respectively to further improve the accuracy of feature description; specifically:
[0100] Low-frequency fractal dimension: mainly reflects the overall complexity of the image, such as large-scale structure, illumination changes and other information;
[0101] High-frequency fractal dimension: mainly reflects local details, such as edges, textures, noise and other complex information.
[0102] In order to further enhance the robustness of the calculation, the present invention can combine the calculation results of different methods and adopt a weighted average method to obtain the final fractal dimension; assuming that the fractal dimensions calculated by different methods are D f1 ,D f2 ,D f3 , the weighted average calculation formula is as follows:
[0103] D f =w1D f1 +w2D f2 +w3D f3 ;
[0104] Among them, D f is the final fractal dimension; D f1 ,D f2 ,D f3 are the fractal dimensions calculated by the box counting method, power spectrum method and multi-scale fractal analysis respectively; w1, w2, w3 are the corresponding method weights, which can be optimized and adjusted according to experimental data.
[0105] Generally speaking, in applications such as image segmentation and pattern recognition, fractal dimensions of multiple scales can be fused to improve the ability to distinguish images of different categories. For example, in the analysis of opera stage scenes, fractal dimensions can be used to distinguish the complexity of different materials, costumes, and background elements, thereby improving the accuracy of image analysis.
[0106] As a possible implementation method, the present invention can calculate the fractal dimension on high-frequency and low-frequency components of different scales respectively, and perform normalization processing in combination with the entropy analysis method; its mathematical formula is as follows:
[0107]
[0108] Where: D′ f is the normalized fractal dimension; max(D f ) and min(D f ) represent the maximum and minimum fractal dimensions calculated at all scales.
[0109] This can eliminate the influence of scale on the calculation results and make the fractal dimension more stable in different scenarios.
[0110] S4. Determine the size and distribution of the convolution kernel according to the fractal dimension, and perform feature extraction;
[0111] In step S3, by calculating the fractal dimension of the image, we quantify the complexity information of the image at different scales; based on this information, we can adaptively adjust the size and distribution of the convolution kernel according to the complexity of the region, thereby achieving efficient feature extraction; this process can select the most appropriate convolution kernel for the complexity characteristics of different regions, so that both the details and the global structure of the image can be fully captured, improving the performance in subsequent processing (such as image classification, recognition or segmentation).
[0112] According to the calculation results of the aforementioned fractal dimension, different regions of the image have different complexities. For example, in regions with high fractal dimensions, the image details are more complex and contain more edge and texture information. In regions with low fractal dimensions, the image structure is smoother and has fewer details. In this case, in order to efficiently extract features, we determine the size of the convolution kernel based on the fractal dimension of the region.
[0113] Generally speaking, areas with higher fractal dimensions indicate richer image details and require smaller convolution kernels for fine feature extraction; while areas with lower fractal dimensions indicate simpler image structures and require larger convolution kernels to capture more macroscopic image information.
[0114] In this embodiment, the size of the convolution kernel can be determined by the following formula:
[0115]
[0116] Where: K is the size of the convolution kernel (usually an integer, representing the side length of the convolution kernel); C is the size of the maximum convolution kernel, such as 9×9; D f is the fractal dimension of the region; D0 is the set fractal dimension threshold, which is often a fixed boundary; k is a constant that controls the rate of change of the convolution kernel size, which is usually adjusted according to experimental data; e is the base of the natural logarithm.
[0117] This formula is obtained by the fractal dimension D f To control the size of the convolution kernel, when the fractal dimension is high, the convolution kernel size is small, and when the fractal dimension is low, the convolution kernel size is large; by dynamically adjusting the size of the convolution kernel, image features can be effectively extracted in areas of different complexity.
[0118] After determining the size of the convolution kernel, we also need to determine the distribution of the convolution kernel based on the complexity of the region; based on the calculation results of the fractal dimension, we can choose different convolution kernel distribution strategies in different regions.
[0119] Specifically, in high fractal dimension areas, due to the complex details and rich edge and texture information, it is necessary to use a smaller convolution kernel and focus on the detailed information of the local area; in order to effectively capture this information, a dense convolution kernel distribution can be selected; a common practice is to use local small convolution kernels, such as 3×3 or 5×5, to perform fine-grained feature extraction in the image.
[0120] In the low fractal dimension area, since the image structure is relatively simple, the main goal is to capture global features or large-scale structures; therefore, a larger convolution kernel can be selected, such as 7×7 or 9×9, to cover a wider area and extract more macroscopic image features.
[0121] The distribution design of convolution kernels is usually optimized in the following ways:
[0122] High fractal dimension area: localized convolution kernel distribution makes the convolution kernel more concentrated in the detail area of the image;
[0123] Low fractal dimension area: The larger scale convolution kernel distribution allows the convolution kernel to cover a larger area of the image and capture the global structure.
[0124] This adaptive convolution kernel distribution design can select the most appropriate convolution kernel size and distribution strategy based on the complexity information of different regions, ensuring that image features can be extracted efficiently and accurately.
[0125] Once the size and distribution of the convolution kernel are determined, the feature extraction process can begin. In this embodiment, the feature extraction process relies on the convolution layer of a convolutional neural network (CNN). In each convolution operation, a convolution kernel that is dynamically adjusted according to the fractal dimension is used. The convolution kernel slides across the image to extract local or global features.
[0126] Through this operation, the network can extract different levels of features from the image layer by layer:
[0127] Local features: such as edges, textures, corners and other detailed information, mainly coming from high fractal dimension areas;
[0128] Global features: such as large-scale structures, regional illumination changes and other information, mainly come from low fractal dimension areas.
[0129] These feature information will be gradually combined to form high-level image feature representations, which can be used for subsequent image analysis tasks such as image classification, target detection, image segmentation, etc.
[0130] In order to further improve the effect of feature extraction, this embodiment can also introduce a multi-scale feature extraction strategy; in this strategy, multiple convolution operations are performed on the image using multiple convolution kernels of different scales to extract feature information at different levels from local to global; a common practice is to use convolution kernels of different sizes such as 3×3, 5×5, 7×7, etc., to capture image information at different scales.
[0131] The advantage of multi-scale feature extraction is that it can capture both details and global information at the same time, avoiding the problem of losing some important features at a single scale; specifically, areas with rich details can be finely extracted using smaller convolution kernels, while areas with simple structures can capture global information using larger convolution kernels, ensuring the comprehensiveness and accuracy of feature extraction.
[0132] In some embodiments, to further improve the accuracy of the model, features extracted at multiple scales may be fused. The fusion operation typically includes:
[0133] Weighted average: Perform weighted average on features extracted at different scales to fuse feature information at different scales;
[0134] Stitching: Features at different scales are stitched together to preserve the detailed information at each scale.
[0135] Through feature fusion, the feature expression ability of the image can be effectively improved, thereby further improving the accuracy of subsequent image processing tasks.
[0136] S5. Calculate the distribution information of the image and adjust the computing resources of each region in the feature extraction process based on the information entropy;
[0137] In the aforementioned steps, the feature information of the image has been preliminarily processed, including adjusting the size and distribution of the convolution kernel based on the fractal dimension, to optimize the effect of feature extraction. However, since the content complexity of the image often has large regional differences, directly using fixed computing resources for feature extraction may result in waste or insufficient computing resources. Therefore, in step S5, information entropy is introduced to calculate the distribution information of the image, and the computing resources in the feature extraction process are adjusted according to the information entropy value of each region, so that the allocation of computing resources is more reasonable.
[0138] In this embodiment, the input image is first locally divided, and the distribution information of each local area is calculated; the local area division method can be a fixed-size grid (such as m×m pixel blocks), or an adaptive division method can be used, such as regional segmentation based on superpixel segmentation or edge detection; after division, the information entropy is calculated in each area to measure the information complexity of the area.
[0139] The calculation formula of information entropy is as follows:
[0140]
[0141] Where: H(x) represents the information entropy of region x; N is the number of different grayscale values in the region; p(x i ) represents the gray value x i The probability of being in this area is calculated as follows:
[0142]
[0143] Where: n(x i ) represents the gray value x i The number of times it appears in the area; M is the total number of pixels in the area.
[0144] The larger the information entropy value, the more complex the grayscale value distribution in the area and the richer the information it contains; conversely, the area with a smaller information entropy indicates that its grayscale value distribution is more uniform and contains less information.
[0145] In this embodiment, after calculating the entropy value of each region, a computing resource allocation strategy based on the entropy value is adopted, so that the high entropy region obtains more computing resources, while the low entropy region is allocated less computing resources; the computing resource allocation formula is as follows:
[0146]
[0147] Among them, R i Assigned to region x i The computing resources represent the computing power share of the region in the feature extraction process; α is the computing resource adjustment coefficient, which controls the allocation ratio of the total computing resources, and its value range is generally 0<α≤1; H(x i ) is the region x i The information entropy of measures the grayscale information complexity of the area and is calculated as follows:
[0148]
[0149] Where M is the area x i The number of different gray values within p(x i,k ) is the gray value x i,k The probability of being in this area is Where n(x i,k ) is the gray value x i , k appears, T is the total number of pixels in the area; N is the total number of image regions, that is, the number of sub-regions into which the entire image is divided; C is the total amount of computing resources available to the entire system, usually referring to the computing power of GPU or CPU, which can be specifically expressed as FLOPS (floating point operations), number of computing threads or computing power budget; H(xj ) is the region x j The information entropy of , where j = 1, 2, ..., N, that is, the sum of the entropy values of all regions is used for normalization calculation, so that computing resources are allocated proportionally.
[0150] In this computing strategy, different computing resource adjustment methods can be used. For example:
[0151] Entropy linear mapping: allocate computing resources directly according to the entropy ratio;
[0152] Threshold segment mapping: preset multiple entropy value intervals, and different intervals use different computing resource levels;
[0153] Exponential decay strategy: The allocation of computing resources is controlled by an exponential function to smooth the adjustment range of computing resources.
[0154] In some embodiments, a dynamic filter adjustment strategy based on entropy values can be adopted; for example, when the entropy value of a region is high, a more complex deep convolutional network (DCNN) is used to extract detailed features, while areas with lower entropy values use lightweight filters or downsampling operations to reduce computational overhead.
[0155] In one possible implementation, the parameters of the feature extraction layer are adjusted based on information entropy; for example:
[0156] Convolution kernel size adjustment: Use larger convolution kernels in high entropy areas to extract more high-level features, and use smaller convolution kernels in low entropy areas to reduce computational complexity;
[0157] Dynamic adjustment of the number of layers: Deeper feature extraction layers can be used in high-entropy areas to enhance information expression capabilities, while the number of feature extraction layers can be reduced in low-entropy areas to improve computational efficiency;
[0158] Computational resolution adjustment: Downsampling is performed in low-entropy areas to reduce the amount of computation, while maintaining high-resolution feature maps in high-entropy areas to ensure complete extraction of detail information.
[0159] In addition, in some embodiments, multi-scale analysis techniques, such as pyramid decomposition or scale-invariant feature transform (SIFT), can be combined to further enhance the adaptability of feature extraction; for example, areas with high entropy values can use high-resolution feature maps for multi-scale feature extraction, while areas with low entropy values can directly skip some scales to reduce the computational burden.
[0160] In some embodiments, the effectiveness of adjusting computing resources using information entropy can be verified through experiments; for example:
[0161] Computing resource utilization: Compare computing resource usage under different entropy adjustment strategies;
[0162] Feature extraction accuracy: Analyze the accuracy of image feature extraction under different entropy adjustment methods, such as the IoU (Intersection over Union) indicator in the object detection task;
[0163] Processing speed: Measure the impact of different entropy computing resource adjustment strategies on computing speed, and observe the computing frame rate (FPS) before and after optimization.
[0164] Through these experimental analyses, we can further optimize the entropy adjustment strategy, such as adjusting the proportion of computing resource allocation, optimizing the computing resource mapping function, or combining the attention mechanism to further improve the effectiveness of feature extraction.
[0165] S6, optimize the extracted features by combining nonlinear dynamics model and use nonlinear activation function to adjust information propagation;
[0166] In the aforementioned steps, the allocation of computing resources has been optimized through information entropy analysis, so that the feature extraction process can adaptively adjust the computing power, thereby improving the efficiency and accuracy of feature extraction; however, due to the highly nonlinear changes in image features during the extraction process, it is difficult to fully model these complex relationships by simply relying on traditional linear transformations. Therefore, in this step S6, the extracted features are further optimized in combination with the nonlinear dynamic model, and the information propagation process is adjusted by a nonlinear activation function to enhance the network's ability to express complex patterns; through this method, the discrimination and stability of feature extraction can be improved while maintaining the utilization of computing resources, ensuring that a more discriminative feature representation is ultimately obtained.
[0167] In this embodiment, the extracted features are optimized based on a nonlinear dynamics model to enhance the ability to capture complex feature patterns; nonlinear dynamics is widely used in complex system modeling, including chaotic systems, bifurcation phenomena and nonlinear transformations, and its application in deep learning can effectively improve feature representation capabilities.
[0168] Specifically, the present invention uses the state evolution equation of the nonlinear dynamic system to optimize the feature representation. The input feature is set as F(t), and the optimized feature representation is G(t). Then its dynamic evolution equation is as follows:
[0169]
[0170] Where: G(t) represents the optimized feature representation; F(t) is the original input feature; β is the feature attenuation coefficient, which controls the dynamic evolution speed of the feature; γ is the feature driving coefficient, which adjusts the influence of the input feature on the final optimization result; f(F(t)) is the nonlinear transformation function of the input feature, usually a nonlinear model such as Logistic mapping, Duffing equation or VanderPol equation is selected; δ is the external noise control term, which is usually set as a random disturbance term to enhance the robustness of the system.
[0171] As an option, nonlinear transformations in the form of logistic mapping can be used to enhance feature expression capabilities. Its mathematical form is as follows:
[0172] f(F(t))=λF(t)(1-F(t));
[0173] Where: λ is a parameter that controls the nonlinear strength, and its value is usually set between 0<λ≤4 to ensure stable dynamic behavior during feature optimization.
[0174] In some embodiments, the Duffing oscillator equation can be used to dynamically optimize the features, which has the following form:
[0175]
[0176] Where: α is the system damping coefficient, which controls the smoothness of the feature optimization process; ω is the system natural frequency, which adjusts the oscillation mode of the feature; κ is the nonlinear term coefficient, which enhances the nonlinear effect of feature optimization.
[0177] This feature optimization method based on nonlinear dynamics can enhance the dynamic adaptability of features, enabling them to better characterize complex patterns while avoiding feature loss caused by over-smoothing.
[0178] After feature optimization, this embodiment further uses a nonlinear activation function to adjust the way information is propagated to enhance the network's expressive power. Traditional linear transformations can only perform simple affine transformations in the feature space, while nonlinear activation functions can introduce highly nonlinear mappings, enabling the model to better learn complex feature patterns.
[0179] Specifically, in an embodiment of the present invention, the choice of activation function is not limited to the traditional ReLU (Rectified Linear Unit), but also combines functions with more nonlinear characteristics, such as Swish, Mish, GELU (Gaussian Error Linear Unit), etc., to enhance the generalization ability of the model.
[0180] As an option, the Swish activation function is used in this embodiment, which has the following form:
[0181] σ(x)=x·sigmoid(βx);
[0182] Among them, σ(x) is the output activation value; x is the input feature; β is a trainable parameter used to adjust the degree of nonlinearity.
[0183] In some embodiments, a Mish activation function may be used, which is expressed as follows:
[0184] σ(x)=x·tanh(log(1+e x ))
[0185] This activation function can provide smooth gradient changes in small gradient areas, enhance the gradient fluidity of the model, and make information propagation more stable.
[0186] In addition, as an option, the present invention can be combined with an adaptive adjustment mechanism of nonlinear activation transformation, that is, the activation function is adaptively selected based on the entropy value calculated by the aforementioned information entropy. For example:
[0187] When the information entropy of a certain area is high, a stronger nonlinear activation function (such as Mish) is used to enhance the learning ability of complex patterns;
[0188] When the information entropy is low, a smoother activation function (such as ReLU) is used to reduce the risk of overfitting.
[0189] In one possible implementation, the activation function can be dynamically selected in conjunction with a gating mechanism:
[0190] Act(x)=w1ReLU(x)+w2Swish(x)+w3Mish(x);
[0191] Among them, w1, w2, and w3 are weight coefficients that can be learned through the training process to optimally combine the characteristics of different activation functions.
[0192] In general, the present invention can combine feature optimization with nonlinear activation function selection through iterative optimization to form an optimization closed loop. For example:
[0193] Phase 1: First, the features are optimized through nonlinear dynamic models to enhance the complex patterns in the feature space;
[0194] The second stage: Use adaptive activation functions to adjust information propagation so that the optimized features can be better propagated in the deep neural network, avoiding gradient vanishing or gradient exploding problems.
[0195] In one possible implementation, the nonlinear dynamics optimization process can be embedded into the batch normalization layer (BatchNormalization) or residual connection (ResidualConnection) of the deep neural network to enhance its ability to regulate information and make the entire optimization process more robust.
[0196] S7, fusing the optimized features and outputting the segmentation result of the opera stage scene;
[0197] In the aforementioned steps, the features of the opera stage scene have been extracted and optimized, and the feature expression ability has been improved through the nonlinear dynamic model. At the same time, a nonlinear activation function has been used to regulate information propagation to enhance the ability to capture complex patterns. However, due to the complexity of the opera stage scene, relying solely on single-scale features for segmentation may result in blurred boundaries or information loss. Therefore, in this step S7, a multi-scale feature fusion strategy is further proposed, and combined with an adaptive weight mechanism and a global-local collaborative optimization method to ensure that the fused features can accurately express the structural information of the foreground and background, and ultimately achieve high-quality opera stage scene segmentation.
[0198] In this embodiment, in order to improve the segmentation accuracy of the opera stage scene, a feature pyramid network (FPN) is combined with adaptive fusion weights to fuse features of different scales to take into account both global scene information and local detail information.
[0199] Feature pyramid fusion, assuming that the optimized multi-scale feature representation is:
[0200] F l ={F1,F2,...,F L};
[0201] Among them, F l Represents the feature representation at different scales l; L is the total number of feature levels, usually selected as 4≤L≤6 to ensure full fusion of multi-scale information.
[0202] In order to perform feature fusion, the high-level features are first upsampled (Upsampl ing) to keep the same spatial resolution as the low-level features. The upsampling process can use bilinear interpolation or learnable transposed convolution (TransposedConvolution), which is calculated as follows:
[0203]
[0204] in, is the feature after upsampling; Upsample(·) is the upsampling operation; F l+1is the l+1th layer feature map, and its spatial resolution is adjusted during the upsampling process; s is the upsampling factor, which is usually set to 2 so that the feature resolution is aligned layer by layer.
[0205] Then, a layer-by-layer fusion mechanism is used to perform weighted fusion of the upsampled features with the features of the current layer:
[0206]
[0207] Among them, F l ′ represents the fused feature map of the first layer; W1, W2 are trainable fusion weights; b is the bias term.
[0208] In some embodiments, an adaptive fusion weight mechanism can be further introduced, that is, according to the importance of features at different levels
[0209] Dynamically adjust the fusion ratio:
[0210]
[0211] Among them, H(F l ) is the characteristic entropy, which is calculated as follows:
[0212]
[0213] Where M is the number of feature channels; p(F l,k ) is the probability of the kth feature channel.
[0214] Through this adaptive weight mechanism, it can be ensured that features with larger information content obtain higher fusion weights, while features with higher information redundancy contribute less, thereby improving the effectiveness of the fusion features.
[0215] After feature fusion is completed, this embodiment further adopts a global-local collaborative optimization strategy to enhance the clarity of the foreground target and suppress the influence of background noise on the segmentation result.
[0216] Foreground extraction based on adaptive threshold. Generally, directly using a fixed threshold for segmentation will result in loss of details. Therefore, this embodiment adopts an adaptive threshold method based on local contrast, which is calculated as follows:
[0217] T(x,y)=μ(x,y)+λ·σ(x,y);
[0218] Where T(x,y) is the adaptive segmentation threshold; μ(x,y) is the local mean; σ(x,y) is the local standard deviation; λ is the adjustment coefficient, which is usually 0.5≤λ≤1.5.
[0219] The segmentation result S(x,y) is determined by the following decision function:
[0220]
[0221] Among them, S(x,y) is the segmentation mask after binarization, 1 represents the foreground (such as opera actors or stage props), and 0 represents the background.
[0222] Based on the boundary optimization of the conditional random field (CRF), in some embodiments, in order to further optimize the segmentation boundary, the conditional random field (CRF) post-processing method is adopted to make the segmentation result smoother in the boundary area and avoid local artifacts. The energy function of CRF is as follows:
[0223] E(S)=∑ i ψ u (S i )+∑ i,j ψ p (S i ,S j );
[0224] Among them, ψ u (S i ) is a single point energy term, which represents the segmentation confidence based on pixel features; ψ p (S i ,S j ) is a two-point energy term used to encourage adjacent pixels to have similar labels and is defined as follows:
[0225]
[0226] Among them, I i ,I j is the color information of pixel i, j; x i ,x j is the pixel coordinate; σ α ,σ β is the smoothing parameter; ω1, ω2 are weighting coefficients.
[0227] Through CRF optimization, the foreground area is ensured to be more coherent while the mis-segmentation of the background area is reduced.
[0228] See also Figure 2 The present invention also provides a drama stage scene segmentation system based on a lightweight backbone structure, comprising:
[0229] The input module is responsible for receiving the opera stage scene images to be segmented and performing preliminary preprocessing. Specifically, the input images are first size-normalized to ensure that they meet the requirements of subsequent feature extraction and processing. The preprocessing process may involve image denoising, brightness and contrast adjustment, etc., to improve the processing effect of subsequent modules. The design of this module takes into account the diversity of images and can automatically adapt to input data of different sources and quality, ensuring that subsequent modules can work efficiently under different conditions.
[0230] The feature extraction module uses wavelet transform technology to extract multi-scale features from the input image. Wavelet transforms can decompose the multi-level and frequency information of an image, extracting rich detailed features. These features are particularly important for describing complex opera stage scenes, which often contain fine textures and complex object shapes. Through multi-scale transformations, this module effectively captures both high-frequency and low-frequency features, providing more comprehensive feature data for subsequent segmentation tasks.
[0231] Furthermore, during feature extraction, the module calculates the fractal dimension of each region as an indicator of image complexity. This reflects the geometric complexity of a local region within an image, helping to assess the level of detail in different regions and thus providing a basis for subsequent resource allocation. By extracting these multi-level features, the module ensures that the key information of the input image is retained to the greatest extent possible.
[0232] The computational resource allocation module dynamically adjusts the allocation of computational resources to different regions based on extracted features, particularly the calculated fractal dimension and information entropy. Information entropy is a measure of image uncertainty and complexity. By analyzing the entropy values of different regions, the module can determine which areas contain more information and which areas may be redundant or contain less background information. Based on this data, the module allocates more computational resources to areas with high information content to ensure accurate segmentation results.
[0233] The core goal of this module is to optimize the computational process, avoid resource waste, and improve segmentation accuracy in highly complex areas. Through this dynamic resource allocation strategy, the system can effectively respond to the processing needs of scenarios of varying complexity. Especially in resource-constrained situations, it can intelligently allocate computing resources and improve overall processing efficiency.
[0234] The information propagation optimization module is a key component of the entire system. Its primary task is to optimize the propagation path of feature information, ensuring that information is effectively adjusted from the input image to the final segmentation result. To address the complexity of opera stage scenes, this module uses a nonlinear dynamics model to optimize information propagation. This model enables the module to handle complex nonlinear relationships in the image and ensures that information is effectively propagated and integrated between feature maps at different levels.
[0235] This module uses nonlinear activation functions to adjust how information is transmitted. Specifically, during the fusion of high-level and low-level features, it regulates the nonlinear relationship between features, enhancing key information while effectively suppressing unimportant background information. This approach not only optimizes information transmission but also enhances the perception of details in complex scenes, ensuring high-quality segmentation under varying lighting and background conditions.
[0236] The fusion and output module is responsible for integrating the optimized information features and outputting the final segmentation result. This module uses multiple feature fusion strategies to weightedly fuse feature information from different levels and scales, ultimately generating a comprehensive feature representation. By fusing information from different sources, the module preserves the unique strengths of each feature while suppressing redundant information.
[0237] When generating the segmentation results, the module combines the fused features with the previously calculated adaptive threshold to ensure precise distinction between foreground and background. Ultimately, the resulting segmentation output clearly identifies key elements on stage, such as actors and props, while avoiding interference from background areas. This process involves more than simple binarization; it involves dynamic adjustments that take into account image content and background complexity, ensuring that the final segmentation output meets high-precision requirements.
[0238] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for segmenting opera stage scenes based on a lightweight backbone structure, characterized in that: The following steps are involved: S1, obtaining an opera stage scene image to be segmented and performing image preprocessing; S2. Decomposing the image by wavelet transform to obtain high-frequency components and low-frequency components of the scale; S3, calculating the fractal dimensions of the high-frequency components and the low-frequency components of each scale of the image, and adjusting subsequent calculation parameters based on the fractal dimensions; S4. Determine the size and distribution of the convolution kernel according to the fractal dimension, and perform feature extraction; S5. Calculate the distribution information of the image and adjust the computing resources of each region in the feature extraction process based on the information entropy; S6, optimize the extracted features by combining nonlinear dynamics model and use nonlinear activation function to adjust information propagation; S7. Fuse the optimized features and output the segmentation results of the opera stage scene.
2. The method for segmenting opera stage scenes based on a lightweight backbone structure according to claim 1, characterized in that: The wavelet transform is decomposed using a Gaussian wavelet function or a Mexican hat wavelet function.
3. The method for segmenting opera stage scenes based on a lightweight backbone structure according to claim 1, characterized in that: The fractal dimension is calculated by a box counting method, which includes covering image regions of different scales and counting the number of minimum covering boxes required to determine the fractal dimension of the region.
4. The method for segmenting opera stage scenes based on a lightweight backbone structure according to claim 1, wherein: The computing resource adjustment is based on information entropy to calculate information complexity, and by setting a threshold, more computing resources are allocated to areas where the information entropy is higher than the threshold.
5. The method for segmenting opera stage scenes based on a lightweight backbone structure according to claim 1, characterized in that: The convolution kernel adopts a variable-scale convolution kernel, and the size and expansion rate of the convolution kernel are dynamically adjusted based on the calculation result of the fractal dimension.
6. The method for segmenting opera stage scenes based on a lightweight backbone structure according to claim 1, characterized in that: The nonlinear dynamics optimization includes using nonlinear dynamics equations to describe the information propagation process and optimizing and adjusting the characteristics based on the information propagation rate.
7. The method for segmenting opera stage scenes based on a lightweight backbone structure according to claim 1, characterized in that: The nonlinear activation function includes a Swish activation function or a LeakyReLU activation function, and the activation function parameters are dynamically adjusted according to feature data of different scales.
8. The method for segmenting opera stage scenes based on a lightweight backbone structure according to claim 1, characterized in that: The feature fusion adopts a weighted feature fusion strategy, and feature information of different scales is fused according to the weight of its fractal dimension, and the weight distribution is calculated based on the contribution of features of each scale.
9. A system for segmenting opera stage scenes based on a lightweight backbone structure, applied to the method for segmenting opera stage scenes based on a lightweight backbone structure according to any one of claims 1 to 8, characterized in that: include: An input module is used to receive the opera stage scene image to be segmented and perform preprocessing; A feature extraction module is used to extract multi-scale features of the input image using wavelet transform and calculate the fractal dimension; Computing resource allocation module, used to adjust the computing resources of each area according to the calculation results of fractal dimension and information entropy; Information propagation optimization module, which is used to optimize information propagation by combining nonlinear dynamics and adjust feature information transmission through nonlinear activation functions; The fusion and output module is used to fuse the optimized feature information and output the segmentation results of the opera stage scene.
10. The opera stage scene segmentation system based on a lightweight backbone structure according to claim 9 is characterized in that: The feature extraction module adopts a dilated convolutional neural network based on wavelet transform to improve the multi-scale capability of feature extraction.