Image segmentation method, device, medium and product

By determining the mutual information graph of visual and audio features through a pre-trained segmentation model, visual features are enhanced to solve the cross-modal alignment problem, achieving efficient and accurate sound source segmentation in complex scenes and improving the accuracy of segmentation masks.

CN122435503APending Publication Date: 2026-07-21SHANGHAI MAIQIXINZHI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI MAIQIXINZHI TECHNOLOGY CO LTD
Filing Date
2026-04-24
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing audiovisual segmentation methods struggle to accurately quantify the fine-grained dependencies between audio and visual modalities in cross-modal alignment, resulting in poor multimodal processing performance in complex scenarios.

Method used

By using a pre-trained segmentation model to determine the visual features of video frame sequences and the audio features of audio data, and by using mutual information graphs to enhance the regions in the visual features that correspond to the audio data, enhanced visual features are generated, thus achieving cross-modal fusion.

Benefits of technology

It improves the accuracy of image segmentation results, especially in achieving accurate and efficient sound source segmentation in complex scenes, and enhances the accuracy of segmentation masks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435503A_ABST
    Figure CN122435503A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an image segmentation method, device, medium and product, and belong to the technical field of image processing. The method comprises: acquiring a video frame sequence and audio data corresponding to the video frame sequence; inputting the video frame sequence and the audio data into a pre-trained segmentation model to obtain a segmentation mask for each video frame in the video frame sequence; wherein the pre-trained segmentation model is used to determine visual features corresponding to the video frame sequence, determine audio features corresponding to the audio data, determine a mutual information graph between the audio features and the visual features, and enhance a region in the visual features that has a corresponding relationship with the audio data based on the mutual information graph to obtain enhanced visual features; and determine the segmentation mask for each video frame in the video frame sequence based on the enhanced visual features. The technical solution of the embodiments of the present application can improve the accuracy of image segmentation in an audio-visual segmentation scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to an image segmentation method, apparatus, medium, and product. Background Technology

[0002] Audio-Visual Segmentation (AVS) is a multimodal learning task that aims to locate and segment objects emitting sound in a video frame sequence by modeling the semantic correspondence between audio and visual signals. This technology has significant application value in fields such as autonomous driving, robotics, augmented reality, and intelligent surveillance, enhancing scene understanding and decision-making capabilities by recognizing sound-emitting objects. However, existing AVS methods still face challenges in cross-modal alignment. This is because current fusion strategies struggle to accurately quantify intermodal correlations.

[0003] To address the aforementioned issues, there is an urgent need for an audiovisual segmentation method that can accurately quantify the fine-grained dependencies between audio and visual modalities, in order to meet the multimodal processing requirements in complex scenarios. Summary of the Invention

[0004] This invention provides an image segmentation method, device, medium, and product that improves the accuracy of image segmentation results.

[0005] According to one aspect of the present invention, an image segmentation method is provided, the method comprising: Obtain the video frame sequence and the corresponding audio data; The video frame sequence and the audio data are input into a pre-trained segmentation model to obtain a segmentation mask for each video frame in the video frame sequence. The pre-trained segmentation model is used to: determine visual features corresponding to the video frame sequence; determine audio features corresponding to the audio data; determine a mutual information graph between the audio features and the visual features; and enhance the regions in the visual features that correspond to the audio data based on the mutual information graph to obtain enhanced visual features; and determine a segmentation mask for each video frame in the video frame sequence based on the enhanced visual features.

[0006] According to another aspect of the present invention, an image segmentation apparatus is provided, the apparatus comprising: Obtain the video frame sequence and the corresponding audio data; The video frame sequence and the audio data are input into a pre-trained segmentation model to obtain a segmentation mask for each video frame in the video frame sequence. The pre-trained segmentation model is used to: determine visual features corresponding to the video frame sequence; determine audio features corresponding to the audio data; determine a mutual information graph between the audio features and the visual features; and enhance the regions in the visual features that correspond to the audio data based on the mutual information graph to obtain enhanced visual features; and determine a segmentation mask for each video frame in the video frame sequence based on the enhanced visual features.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When one or more programs are executed by one or more processors, the one or more processors implement any of the image segmentation methods described in the embodiments of this disclosure.

[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute any of the image segmentation methods of the present invention.

[0009] According to another aspect of the present disclosure, a computer program product is provided, which, when executed by a processor, implements any of the image segmentation methods described in the embodiments of the present disclosure.

[0010] The technical solution provided by this invention involves a pre-trained segmentation model that determines visual features corresponding to a video frame sequence and audio features corresponding to audio data. It also determines a mutual information graph between the audio and visual features, and enhances the regions in the visual features that correspond to the audio data based on the mutual information graph, thus obtaining enhanced visual features. This transforms the semantic information of the audio into spatial attention guidance, allowing the visual feature space to focus on the sound-producing object, ensuring the accuracy of the enhanced visual features. This enables the enhanced visual features to achieve accurate and efficient sound source segmentation in complex scenes, thereby improving the accuracy of the segmentation mask determined based on the enhanced visual features.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A schematic flowchart of the image segmentation method provided in the embodiments of this disclosure; Figure 2 This is yet another flowchart illustrating the image segmentation method provided in this embodiment of the disclosure; Figure 3 A schematic flowchart illustrating the model training method provided in this embodiment of the disclosure; Figure 4 This is a schematic diagram of the structure of the image segmentation apparatus provided in the embodiments of this disclosure; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0016] Figure 1This is a schematic diagram of an image segmentation process provided in an embodiment of this disclosure. This embodiment is applicable to pixel-level segmentation tasks of sound objects in videos. The method can be executed by an image segmentation device, which can be implemented in hardware and / or software and can be configured in electronic devices such as computers or servers. Figure 1 As shown, the method in this embodiment includes: S110. Obtain the video frame sequence and the audio data corresponding to the video frame sequence.

[0017] A video frame sequence refers to video captured by a camera. The corresponding audio data refers to the audio data synchronously acquired by an audio acquisition device during the acquisition of the video frame sequence.

[0018] In one embodiment, the visual frame sequence and its corresponding audio data can be selected as a video frame sequence and audio data within a predetermined time window. The width of the predetermined time window can be configured.

[0019] In scenarios such as autonomous driving, robotics, augmented reality, and intelligent monitoring, video frame sequences and corresponding audio data are collected simultaneously.

[0020] S120. Input the video frame sequence and audio data into the pre-trained segmentation model to obtain the segmentation mask for each video frame in the video frame sequence; wherein, the pre-trained segmentation model is used to: determine the visual features corresponding to the video frame sequence; determine the audio features corresponding to the audio data; determine the mutual information graph between the audio features and the visual features; and enhance the regions in the visual features that correspond to the audio data based on the mutual information graph to obtain enhanced visual features; and determine the segmentation mask for each video frame in the video frame sequence based on the enhanced visual features.

[0021] A segmentation mask, also known as a segmentation template, is a binary image. Once the segmentation mask is determined, the corresponding video frames can be segmented according to the segmentation mask to obtain the segmentation results.

[0022] The pre-trained segmentation model includes an encoder, a neural network discriminator, and a decoder.

[0023] In one embodiment, the encoder includes two existing sub-encoders, which are used to determine visual features corresponding to a video frame sequence and audio features corresponding to audio data, respectively.

[0024] In one embodiment, the encoder includes a visual encoder and an audio encoder. The visual encoder may be a pyramid visual encoder (a PVTv2-based pulse visual encoder), and the audio encoder may be a VGGish-based pulse audio encoder. The pyramid visual encoder is used to convert video sequences into pulse-like visual features, providing accurate visual cues for cross-modal sound source localization. The VGGish-based pulse audio encoder is used to convert the Mel spectrogram of audio data into pulse-like audio features, providing accurate auditory cues for visual queries. The Mel spectrogram is the short-time Fourier transform result of the audio data.

[0025] Specifically, the encoder in the pre-trained segmentation model is based on a spatial reduction attention mechanism, which reduces the spatial dimensions of the key matrix K and value matrix V of the visual features to their original dimensions through downsampling. The formula for calculating attention weights is as follows: .

[0026] Where Q is the query projection of the initial resolution feature map of the video frame, K is the key projection of the downsampled feature map, V is the value projection of the downsampled feature map, d is the feature dimension, and R is the reduction ratio, where R is greater than 1. The visual features extracted by this pyramid video encoder can be represented as follows: Each layer of visual features can be represented as Then, based on pulse-based feature activation, visual features are converted into pulse-like visual features, specifically through IF-SR neurons (pulse response neurons): ; ; .

[0027] in, X[t] is the input current (visual feature) at time step t. The membrane potential at time step t represents the internal state of a impulse-response neuron. The membrane potential is identified as time step t-1. For indicator functions, This refers to the output pulse of a spurious response neuron, also known as a spurious visual feature. It is a binary signal indicating whether the spurious response neuron fires a pulse at time step t. Spurious visual features enable sparse, event-driven processing, reducing computational costs.

[0028] First, the audio signal Perform a short-time Fourier transform (STFT) to generate a Mel spectrogram. Synchronized with the video frame time; the audio encoder in the pre-trained segmentation model extracts features from the Mel spectrogram to obtain audio features ( Then, based on pulse-based feature activation, the audio features are converted into pulse-like audio features, as detailed in the aforementioned data processing flow for pulse-response neurons. This step ensures time alignment between pulse-like audio features and pulse-like visual features while reducing data processing power consumption. The pulse-like audio features are also binary pulse sequences to enable efficient event-driven computation.

[0029] The use of pulsed audio features and pulsed visual features can significantly reduce computational complexity and improve computational speed, making it particularly suitable for audiovisual segmentation tasks in low-power edge computing scenarios.

[0030] In one embodiment, the pre-trained segmentation model further includes a neural network discriminator; the neural network discriminator is used to: process the pulsed audio features and the visual features to obtain the mutual information graph (see...). Figure 2 The mutual information graph includes the mutual information value between the audio feature and the spatial position of each video frame in the image sequence; the normalization result of the mutual information graph is determined, and the element-wise multiplication result of the normalization result and the visual feature is used as the enhanced visual feature.

[0031] Specifically, based on a mutual information neural estimation framework. Let the pulsed audio features be... Pulse-like visual features are Where B is the batch size, and C=256. , Flattening pulsed visual features into a set of spatial tokens Based on a two-layer MLP discriminator, the mutual information value between the sound impulse frequency features and each visual spatial location f_V^{(i,j)} is estimated using the following formula: ; Where P is the joint distribution of positive samples, This is the product of the marginal distributions of the negative samples. After all mutual information values ​​are determined, they are summarized into a mutual information graph. The mutual information graph is then normalized to obtain the normalized mutual information graph.

[0032] The specific formula is as follows: ; Then, visual features are enhanced through element-wise multiplication: ; This enhanced visual feature can highlight visual regions that are relevant to audio semantics.

[0033] In one embodiment, the pre-trained segmentation model further includes a pulse Transformer, which comprises an audio query generator, a pulse Transformer decoder, and a segmentation decoder. The audio query generator generates a target audio query vector based on the pulsed audio features. The pulse Transformer decoder fuses the target audio query vector with the enhanced visual features to generate a query result. The segmentation decoder determines the segmentation mask based on the query result and the enhanced visual features. By combining the audio query generator and the pulse Transformer decoder, accurate query results can be obtained, thereby improving the accuracy of the segmentation mask determined by the segmentation decoder.

[0034] Based on the aforementioned example, the audio query generator can be configured to: generate an initial audio query vector corresponding to the pulsed audio features; determine the audio key vector and audio value vector corresponding to the initial audio query vector; fuse the initial audio query vector, the audio key vector, the audio value vector, and a first predetermined scaling factor to obtain an intermediate audio query vector; and inject learnable pulse position information into the intermediate audio query vector to obtain the target audio query vector.

[0035] Specifically, the audio query generator uses pulsed audio features Pulse generation condition query ,in The query quantity is [number]. An initial query vector corresponding to the pulsed audio features is determined based on linear projection, and then the target audio query vector is generated using the following formula: ; ; ; ; .

[0036] in, For integral firing neurons with soft reset, This is the initial query vector. This is a pulse-type audio query vector. It has pulse-like audio characteristics. It is an audio key vector; It is an audio value vector; The first predetermined scaling factor, Expand the channel to Dimensionality is added to enhance feature representation capabilities and provide richer value representations for attention mechanisms. Used to compress the expanded channel dimensions back to the original dimensions, and When used in pairs, they constitute a complete extended-compression strategy. This is the intermediate audio query vector. For the target audio query vector, For learnable pulse position embedding, This is spike-based feature activation.

[0037] The audio query generator refines the coarse initial query vector into a high-quality target audio query vector rich in audio context information through a pulse self-attention mechanism, providing accurate audio-driven queries for subsequent cross-modal fusion.

[0038] Based on the foregoing embodiments, the pulse Transformer decoder is used to: determine the visual key vector and visual value vector corresponding to the flattening result of the enhanced visual features; and fuse the target audio query vector, the visual key vector, the visual value vector, and a second predetermined scaling factor to obtain the query result.

[0039] Specifically, the Transformer decoder fuses the target audio query vector with impulse visual features and generates query results based on a multi-head self-attention mechanism. This is achieved through the following formula: ; ; ; in, The result is a flattening of pulsed visual features. For visual key vectors, For the visual value vector, the second predetermined scaling factor, The results are from the query.

[0040] The Pulse Transformer decoder uses a pulse cross-attention mechanism to enable cross-modal interaction between the target audio query vector and pulsed visual features, allowing the audio query to acquire relevant visual contextual information and generate visually enhanced query results. For example, it matches "heard sounds" with "seen objects".

[0041] Building upon the aforementioned embodiments, the Pulse Transformer comprises three layers, each with eight attention heads and a dimension of 256. The segmentation decoder generates a segmentation mask based on the target audio query vector and the query results. Specifically: ; ; in, For linear transformation, Used to extract the query vector for each target audio. This is the aggregated result of the target audio query vector. To enhance visual features, Used to add height and width to the aggregation result, thereby expanding it into a four-dimensional tensor to match the dimensions of the enhanced visual features; For segmentation mask.

[0042] The segmentation decoder performs the above processing through two layers of convolutional LIF neurons (leaking integral firing neurons, 3×3 nuclei) and three layers of spiking MLP (multilayer perceptron, hidden dimension 512). The target audio query vectors are aggregated into a unified semantic vector, then converted into a spatially fusionable format, and injected into enhanced visual features to generate a cross-modal fusion feature representation, resulting in a segmentation mask. ,in, This represents the number of segmentation categories. Using this segmentation mask to segment the corresponding video frames yields accurate audiovisual segmentation results.

[0043] In one embodiment, the segmentation model is an artificial neural network. After the enhanced visual features are determined, the segmentation mask is determined based on these enhanced visual features; or the segmentation mask is determined based on both enhanced visual features and visual features. Taking the latter as an example, the decoder introduces low-level, high-resolution visual features from the encoder through "skip connections." These low-level features contain rich edge and texture information. During fusion, the decoder combines the low-level visual features with the enhanced visual features of the current layer. The final layer of the decoder is a pixel-wise classifier (e.g., a 1x1 convolutional layer followed by a Softmax or Sigmoid activation function). It takes the high-resolution feature map output by the decoder as input and generates a probability value for each pixel, representing the confidence that the pixel belongs to the "sounding object" category. Finally, a binary map or probability map of the same size as the input image is generated; this is the segmentation mask.

[0044] In summary, the segmentation model in this embodiment can be various forms of machine learning models, such as artificial neural network models and spiking neural network models; and it is foreseeable that a combination model of artificial neural networks and spiking neural networks is also possible, as are other forms of network models. Regarding the combination model of artificial neural networks and spiking neural networks, this combination model includes a first part and a second part. The first part is implemented based on an artificial neural network and may optionally include a video encoder, while the second part is implemented based on a spiking neural network.

[0045] The technical solution provided by this invention involves a pre-trained segmentation model that determines visual features corresponding to a video frame sequence and audio features corresponding to audio data. It also determines a mutual information graph between the audio and visual features, and enhances the regions in the visual features that correspond to the audio data based on the mutual information graph, thus obtaining enhanced visual features. This transforms the semantic information of the audio into spatial attention guidance, allowing the visual feature space to focus on the sound-producing object, ensuring the accuracy of the enhanced visual features. This enables the enhanced visual features to achieve accurate and efficient sound source segmentation in complex scenes, thereby improving the accuracy of the segmentation mask determined based on the enhanced visual features, audio features, and visual features.

[0046] The segmentation model was implemented in PyTorch 2.2, trained for 50 epochs using six NVIDIA A100 GPUs (80GB each) with a batch size of 8. The optimizer used was Adam (Adaptive Moment Estimation Optimizer), with an initial learning rate of 1e-4, decaying to 1e-6 using cosine annealing. After training, the model was tested on the AVSBench dataset, which includes three subtasks: single-source, multi-source, and semantic segmentation, covering approximately 7000 videos and 80,000 frames of pixel-level labeled video data. Each video was typically 5 seconds long with a frame rate of 1 frame / second. Five frames were extracted as visual input, and the audio signal was acquired at a 16kHz mono sampling rate. The video frame sequence was extracted from the video, denoted as... Where T=5, H and W are the frame height and width, respectively. The frame sequence is adjusted to a uniform resolution (224×224) through standard normalization.

[0047] The mean Intersection over Union (mIoU) and F-score were used as the main metrics. Tests were conducted on an NVIDIA A100 GPU with a batch size of 2. The SNN (Spiking Neural Network) was executed in an event-driven manner. Energy consumption was estimated based on operation counts, with ANN calculated at 4.6 pJ per multiply-accumulate (MAC) and SNN at 0.9 pJ per accumulation (AC). When the segmentation model is an artificial neural network, the average intersection-union (IU) accuracy of the segmentation mask is improved by at least 1.8% compared to existing technologies. When the segmentation model is a spiking neural network (SNN), the required energy consumption is reduced by approximately 80.4% compared to existing technologies while maintaining high segmentation accuracy, making it particularly suitable for efficient multimodal processing in low-power edge computing scenarios. When the visual encoder in the segmentation model is an artificial neural network and the other parts are spiking neural networks, the average IU of the segmentation mask is improved by approximately 2.3% compared to existing technologies, while energy consumption is reduced by approximately 74.3%, achieving an excellent balance between accuracy and efficiency. The event-driven nature of spiking neural networks makes them compatible with various hardware forms (such as Intel Loihi, IBM TrueNorth, and Lynxi HE200), suitable for low-power real-time applications such as robots and monitoring equipment. It is understood that modular design can support expansion to more complex multimodal tasks, such as spatiotemporal sound event tracking or interactive audio-visual reasoning, providing a foundation for the development of future multimodal intelligent systems. In summary, this invention provides an efficient and accurate audiovisual segmentation solution by optimizing cross-modal fusion through mutual information, offering a new technical path for multimodal segmentation tasks in edge computing environments.

[0048] The segmentation model can be trained directly on the target sample set to obtain a pre-trained segmentation model; or it can be trained sequentially on multiple sample sets according to a predetermined order.

[0049] Figure 3 This is a flowchart illustrating a model training method provided in an embodiment of the present invention. This embodiment trains a segmentation model sequentially based on multiple sample sets according to a predetermined order. Figure 3 As shown, the method includes: S210. Train the audio encoder and visual encoder based on the first sample set and the second sample set respectively to obtain the intermediate segmentation model, which includes the pre-trained audio encoder and the pre-trained visual encoder.

[0050] The first sample set can be the ImageNet-1K dataset, a publicly available image dataset containing manually annotated images of various objects. The second sample set can be the AudioSet dataset, a publicly available audio dataset containing annotations for multiple types of sound events.

[0051] During the training of an audio encoder, after the audio features are determined, to achieve compatibility with spiking neural networks, a pulse-emission approximation strategy, weight mapping, and activation functions are used to convert continuous activations into a sparse pulse-like audio feature sequence. Similarly, during the training of a visual encoder, after the visual features are determined, a pulse-emission approximation strategy, weight mapping, and activation functions are used to convert continuous activations into a sparse pulse-like visual feature sequence. The implementation process of the pulse-emission approximation strategy is similar to the pulse feature activation corresponding to the aforementioned pulse response neurons, the difference being that during training, the pulse response neurons generate integer pulse counts. Where D=4 is the maximum pulse count. It is a pruning function; during inference, it is expanded into a binary pulse sequence.

[0052] S210. Train the intermediate segmentation model based on the third sample set to obtain the pre-trained segmentation model.

[0053] The third sample set includes video frame sequences, audio data, and segmentation masks for each video frame in the video frame sequence.

[0054] In this embodiment, a composite loss function is preferably used to train the segmentation model. Specifically, the composite loss function is a weighted sum of the intersection-union loss and the mutual information loss; the intersection-union loss is used to minimize the difference between the predicted segmentation mask and the real mask; and the mutual information loss is used to enhance the alignment between audiovisual feature pairs.

[0055] The composite loss function can be expressed as: ; in, Optionally, the cross-union ratio loss based on Dice loss can be used to supervise the consistency between the predicted segmentation mask and the real segmentation mask; Mutual information loss is used to optimize the alignment of audio and visual features; This is a weighting factor used to balance the contributions between the two types of losses. Possible values ​​include 0.1 and 0.15.

[0056] The mutual information loss can be expressed as: ; in, It is a learnable neural discriminator. It is an audio feature. For positive sample visual features, P represents the visual features of negative samples, and P represents the joint distribution of positive samples. E represents the expected value, which is the product of the marginal distributions of the negative samples. Positive and negative sample pairs are constructed based on the true segmentation mask. During training, positive and negative samples are compared and learned using mutual information loss to maximize the mutual information between pulsed audio features and pulsed visual features, optimize the Donsker-Varadhan lower bound loss, and achieve cross-modal alignment.

[0057] Figure 4 This is a schematic diagram of the structure of the image segmentation apparatus provided in an embodiment of this disclosure. Figure 4 As shown, the image segmentation device includes: Data acquisition module 110 is used to acquire video frame sequences and audio data corresponding to the video frame sequences; The data processing module 120 is used to input the video frame sequence and the audio data into a pre-trained segmentation model to obtain a segmentation mask for each video frame in the video frame sequence. The pre-trained segmentation model is used to: determine visual features corresponding to the video frame sequence; determine audio features corresponding to the audio data; determine a mutual information graph between the audio features and the visual features; and enhance the regions in the visual features that correspond to the audio data based on the mutual information graph to obtain enhanced visual features; and determine a segmentation mask for each video frame in the video frame sequence based on the enhanced visual features.

[0058] The technical solution provided by this invention involves a pre-trained segmentation model that determines visual features corresponding to a video frame sequence and audio features corresponding to audio data. It also determines a mutual information graph between the audio and visual features, and enhances the regions in the visual features that correspond to the audio data based on the mutual information graph, thus obtaining enhanced visual features. This transforms the semantic information of the audio into spatial attention guidance, allowing the visual feature space to focus on the sound-producing object, ensuring the accuracy of the enhanced visual features. This enables the enhanced visual features to achieve accurate and efficient sound source segmentation in complex scenes, thereby improving the accuracy of the segmentation mask determined based on the enhanced visual features.

[0059] In one embodiment, the pre-trained segmentation model further includes a pulse pyramid-based visual encoder and a VGGish-based audio encoder. The pulse pyramid-based visual encoder is used to determine the pulsed visual features corresponding to the video frame sequence. The VGGish-based audio encoder is used to determine the pulsed audio features corresponding to the Mel spectrogram, which is the short-time Fourier transform result of the audio data.

[0060] In one embodiment, the pre-trained segmentation model further includes a neural network discriminator, which is used to: The pulsed audio features and the pulsed visual features are processed to obtain the mutual information map, which includes the mutual information values ​​between the pulsed audio features and the spatial positions of each video frame in the video frame sequence. Determine the normalization result of the mutual information graph, and use the result of element-wise multiplication of the normalization result with the visual feature as the enhanced visual feature.

[0061] In one embodiment, the pre-trained segmentation model further includes a pulse Transformer, which includes an audio query generator, a pulse Transformer decoder, and a segmentation decoder; The audio query generator is used to generate a target audio query vector based on the pulsed audio features. The pulse Transformer decoder is used to fuse the target audio query vector with the enhanced visual features to generate query results; The segmentation decoder is used to determine the segmentation mask based on the query result and the enhanced visual features.

[0062] In one embodiment, the audio query generator is used for: The initial query vector corresponding to the pulsed audio features is determined based on linear projection; Determine the pulsed audio query vector corresponding to the initial query vector; Determine the audio key vector and audio value vector corresponding to the pulsed audio features; The pulse-type audio query vector, the audio key vector, the audio value vector, and the first predetermined scaling factor are fused to obtain an intermediate audio query vector; Learnable pulse position information is injected into the intermediate audio query vector to obtain the target audio query vector; The pulse Transformer decoder is used for: Determine the visual key vector and visual value vector corresponding to the flattening result of the enhanced visual features; The target audio query vector, the visual key vector, the visual value vector, and the second predetermined scaling factor are fused to obtain the query result.

[0063] In one embodiment, the segmentation model training process employs a composite loss function, which is a weighted sum of the intersection-union loss and the mutual information loss; The cross-union ratio loss is used to minimize the difference between the predicted segmentation mask and the true mask; The mutual information loss is used to enhance the alignment between audiovisual feature pairs.

[0064] In one embodiment, the segmentation model includes a first part and a second part, the first part being implemented based on an artificial neural network and including a video encoder, and the second part being implemented based on a spiking neural network.

[0065] The image segmentation method apparatus provided in this disclosure can execute the image segmentation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0066] It is worth noting that the various units and modules included in the above-mentioned image segmentation method apparatus are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0067] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0068] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0069] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0070] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as image segmentation methods.

[0071] In some embodiments, the image segmentation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via read-only memory (ROM) 12 and / or communication unit 19. When the computer program is loaded into random access memory (RAM) 13 and executed by processor 11, one or more steps of the image segmentation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the image segmentation method by any other suitable means (e.g., by means of firmware).

[0072] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0073] Computer programs for implementing the image segmentation method of this disclosure may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0074] This disclosure provides a computer-readable storage medium storing computer instructions for causing a processor to execute an image segmentation method, including: Obtain the video frame sequence and the corresponding audio data; The video frame sequence and the audio data are input into a pre-trained segmentation model to obtain a segmentation mask for each video frame in the video frame sequence. The pre-trained segmentation model is used to: determine visual features corresponding to the video frame sequence; determine audio features corresponding to the audio data; determine a mutual information graph between the audio features and the visual features; and enhance the regions in the visual features that correspond to the audio data based on the mutual information graph to obtain enhanced visual features; and determine a segmentation mask for each video frame in the video frame sequence based on the enhanced visual features.

[0075] In the context of this disclosure, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0076] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0077] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0078] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0079] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from ROM 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of embodiments of this disclosure.

[0080] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements an image segmentation method according to any embodiment of this disclosure.

[0081] In implementing a computer program product, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0082] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.

[0083] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An image segmentation method, characterized in that, The method includes: Obtain the video frame sequence and the corresponding audio data; The video frame sequence and the audio data are input into a pre-trained segmentation model to obtain a segmentation mask for each video frame in the video frame sequence. The pre-trained segmentation model is used to: determine visual features corresponding to the video frame sequence; determine audio features corresponding to the audio data; determine a mutual information graph between the audio features and the visual features; and enhance the regions in the visual features that correspond to the audio data based on the mutual information graph to obtain enhanced visual features; and determine a segmentation mask for each video frame in the video frame sequence based on the enhanced visual features.

2. The method according to claim 1, characterized in that, The pre-trained segmentation model also includes a visual encoder based on a pulse pyramid and an audio encoder based on VGGish. The pulse pyramid-based visual encoder is used to determine the pulsed visual features corresponding to the video frame sequence. The VGGish-based audio encoder is used to determine the pulsed audio features corresponding to the Mel spectrogram, which is the short-time Fourier transform result of the audio data.

3. The method according to claim 2, characterized in that, The pre-trained segmentation model further includes a neural network discriminator, which is used for: The pulsed audio features and the pulsed visual features are processed to obtain the mutual information map, which includes the mutual information values ​​between the pulsed audio features and the spatial positions of each video frame in the video frame sequence. Determine the normalization result of the mutual information graph, and use the result of element-wise multiplication of the normalization result with the visual feature as the enhanced visual feature.

4. The method according to claim 2, characterized in that, The pre-trained segmentation model also includes a pulse Transformer, which comprises an audio query generator, a pulse Transformer decoder, and a segmentation decoder. The audio query generator is used to generate a target audio query vector based on the pulsed audio features. The pulse Transformer decoder is used to fuse the target audio query vector with the enhanced visual features to generate query results; The segmentation decoder is used to determine the segmentation mask based on the query result and the enhanced visual features.

5. The method according to claim 4, characterized in that, The audio query generator is used for: The initial query vector corresponding to the pulsed audio features is determined based on linear projection; Determine the pulsed audio query vector corresponding to the initial query vector; Determine the audio key vector and audio value vector corresponding to the pulsed audio features; The pulse-type audio query vector, the audio key vector, the audio value vector, and the first predetermined scaling factor are fused to obtain an intermediate audio query vector; Learnable pulse position information is injected into the intermediate audio query vector to obtain the target audio query vector; The pulse Transformer decoder is used for: Determine the visual key vector and visual value vector corresponding to the flattening result of the enhanced visual features; The target audio query vector, the visual key vector, the visual value vector, and the second predetermined scaling factor are fused to obtain the query result.

6. The method according to claim 1, characterized in that, The segmentation model training process uses a composite loss function, which is a weighted sum of the intersection-union loss and the mutual information loss. The cross-union ratio loss is used to minimize the difference between the predicted segmentation mask and the true mask; The mutual information loss is used to enhance the alignment between audiovisual feature pairs.

7. The method according to claim 1, characterized in that, The segmentation model includes a first part and a second part. The first part is based on an artificial neural network and includes a video encoder, while the second part is based on a spiking neural network.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the image segmentation method as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the image segmentation method according to any one of claims 1-7.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the image segmentation method according to any one of claims 1-7.