A scene target detection method and device based on a visual saliency model

CN118657931BActive Publication Date: 2026-09-25XIAN LINGKONG ELECTRONICS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411067902.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2026-09-25
Estimated Expiration
2044-08-06

AI Technical Summary

Technical Problem

[0005]在本申请实施例中,通过提供一种基于视觉显著性模型的场景目标检测方法,解决了在复杂多变的模拟训练战场的场景中,传统算法往往存在误检率高、鲁棒性差,特别是在目标被遮挡、背景与目标相似度高、光照条件变化大等情况下,传统算法的检测性能会大幅下降,导致传统显著性检测算法在复杂图像中显著值低,不能适应不同的环境和场景变化的问题

Benefits of technology

本申请实施例提供了一种基于视觉显著性模型的场景目标检测方法,通过明确检测目标和设定约束条件,使得模型在训练过程中更加聚焦于场景中的关键目标。有助于提升模型对关键目标的识别精度,减少误检和漏检的情况,从而提高整体检测的准确性。利用包含丰富图像数据及其精确位置和类别标注的数据集进行训练,使得模型能够学习到更多样化的特征表示,使其在不同场景和复杂环境下依然能够保持较高的检测性能。在卷积神经网络和多层感知机之间引入残差块,通过残差连接有效缓解了深层网络训练中的梯度消失或爆炸问题,不仅加快了模型的训练速度,还使得模型在推理阶段能够更快地处理图像数据,提高了检测效率。通过多层感知机的输出层区分图像的显著区域和非显著区域,使得模型能够自动学习并识别出图像中最具代表性的区域,为后续的目标检测提供了更加准确的先验信息,进一步优化了目标检测的效果。构建的模型框架具有良好的灵活性和可扩展性。可以根据实际需求调整残差块的数量、卷积层的配置以及多层感知机的结构,以适应不同复杂度和规模的目标检测任务。解决了在复杂多变的模拟训练战场的场景中,传统算法往往存在误检率高、鲁棒性差,特别是在目标被遮挡、背景与目标相似度高、光照条件变化大等情况下,传统算法的检测性能会大幅下降,导致传统显著性检测算法在复杂图像中显著值低,不能适应不同的环境和场景变化的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118657931B_ABST
    Figure CN118657931B_ABST
Patent Text Reader

Abstract

The application discloses a scene target detection method and device based on a visual saliency model, relates to the technical field of image processing, and comprises the following steps: collecting images in a scene, labeling the positions and categories of targets in the images, and constructing a data set; defining a detection target and setting a constraint condition; constructing a residual block between a convolutional neural network and a multilayer perceptron according to the detection target and the constraint condition, so as to introduce a residual connection, and constructing an initial saliency model; acquiring output data of the residual block as input data of the multilayer perceptron; outputting a first prediction result of the initial saliency model through an output layer of the multilayer perceptron; training the initial saliency model by using the data set, evaluating and optimizing the performance of the initial saliency model on a verification set, and thus obtaining a final saliency model. The problems of high false detection rate and poor robustness of traditional algorithms in a complex and changeable simulated training battlefield scene are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method and apparatus for scene target detection based on a visual saliency model. Background Technology

[0002] In the complex and ever-changing simulated training battlefield, the efficient and accurate extraction and identification of key targets in scene images has become an important foundation for intelligent equipment.

[0003] Traditional saliency detection algorithms primarily rely on low-level features of images, such as color, brightness, and texture, to detect salient regions by simulating the attention mechanism of the human visual system. The simulated training battlefield is a dynamically changing environment where the position, shape, and even state of the target can change at any time, requiring algorithms to possess high real-time performance and adaptability.

[0004] However, in the complex and ever-changing scenarios of simulated training battlefields, traditional algorithms often suffer from problems such as high false detection rates and poor robustness. In particular, when the target is occluded, the background and target are highly similar, or the lighting conditions vary greatly, the detection performance of traditional algorithms will drop significantly, resulting in low saliency values ​​for traditional saliency detection algorithms in complex images, making them unable to adapt to different environments and scene changes. Summary of the Invention

[0005] In this application embodiment, a scene target detection method based on a visual saliency model is provided, which solves the problem that traditional algorithms often have high false detection rates and poor robustness in complex and ever-changing simulated training battlefield scenarios. In particular, the detection performance of traditional algorithms will drop significantly when the target is occluded, the background and the target are highly similar, and the lighting conditions vary greatly. This results in traditional saliency detection algorithms having low saliency values ​​in complex images and being unable to adapt to different environments and scene changes.

[0006] In a first aspect, embodiments of this application provide a scene target detection method based on a visual saliency model. The method includes: collecting images from a scene; labeling the location and category of targets in the images to construct a dataset; defining detection targets and setting constraints; constructing residual blocks between a convolutional neural network and a multilayer perceptron based on the detection targets and constraints to introduce residual connections and construct an initial saliency model; obtaining the output data of the residual blocks as input data to the multilayer perceptron; outputting a first prediction result of the initial saliency model through the output layer of the multilayer perceptron; wherein the first prediction result is used to distinguish between salient and non-salient regions of the image; training the initial saliency model using the dataset; evaluating and optimizing the performance of the initial saliency model on a validation set to obtain a final saliency model.

[0007] In one possible implementation, constructing a residual block between the convolutional neural network and the multilayer perceptron to introduce residual connections, based on the detection target and constraints, includes: defining a main path function and a shortcut function; adding the output of the main path function to the output of the shortcut function to obtain the output data of the residual block, thereby introducing residual connections.

[0008] In one possible implementation, the main path function is used to receive input data and output a feature map after processing the input data through multiple operations; the multiple operations include convolution operations, batch normalization operations, and nonlinear activation function operations.

[0009] In one possible implementation, the input data of the shortcut function is the same as the input data of the main path function; a weight matrix is ​​introduced, and the input data is multiplied by the weight matrix to obtain the output of the shortcut function.

[0010] One possible implementation also includes introducing a multi-scale feature fusion step into the residual block.

[0011] In one possible implementation, the multi-scale feature fusion step includes: performing convolution operations on the input data using multiple convolution kernels of different scales in the main path function to generate feature maps of multiple scales; performing non-uniform segmentation on each generated feature map to divide it into multiple regions of interest of different scales and shapes, so that each feature map has multiple regions of interest; determining the relative position of each region of interest on the same scale feature map and mapping it to its respective feature map; extracting a small region around the center point of each region of interest on the feature map as the mapped feature map; performing convolution operations on each mapped feature map to generate new feature maps with the same number of regions of interest; performing a fusion operation on multiple new feature maps to generate the final feature map; using the final feature map as input data for a multilayer perceptron; and outputting the second prediction result of the initial saliency model through the output layer of the multilayer perceptron; wherein the second prediction result is used to distinguish between salient and non-salient regions of the image.

[0012] In one possible implementation, evaluating and optimizing the performance of the initial saliency model on the validation set to obtain the final saliency model includes: designing evaluation metrics and constructing representative scenarios as simulation scenarios in an equipment intelligent algorithm experimental tool; wherein the evaluation metrics include detection accuracy, false positive rate, false negative rate, and processing speed, and the representative scenarios include scenarios that can cover different environmental conditions, different target types, and different detection difficulties; loading the simulation scenarios and the initial saliency model in a virtual independent operating environment, conducting simulation experiments, and collecting and storing experimental data in real time; after the experiment, evaluating and analyzing the collected experimental data, and adjusting and optimizing the initial saliency model based on the evaluation and analysis results to obtain the final saliency model.

[0013] Secondly, embodiments of this application provide a scene target detection device based on a visual saliency model. The device includes: a collection module for collecting images from a scene, labeling the location and category of targets in the images, and constructing a dataset; a definition module for defining detection targets and setting constraints; a construction module for constructing residual blocks between a convolutional neural network and a multilayer perceptron based on the detection targets and constraints, introducing residual connections, and constructing an initial saliency model; an acquisition module for acquiring the output data of the residual blocks as input data for the multilayer perceptron; an output module for outputting a first prediction result of the initial saliency model through the output layer of the multilayer perceptron; wherein the first prediction result is used to distinguish between salient and non-salient regions of the image; and an optimization module for training the initial saliency model using the dataset, evaluating and optimizing the performance of the initial saliency model on a validation set, thereby obtaining a final saliency model.

[0014] Thirdly, embodiments of this application provide a scene target detection server based on a visual saliency model, including a memory and a processor; the memory is used to store computer-executable instructions; the processor is used to execute the computer-executable instructions to implement the method described in the first aspect or any possible implementation of the first aspect.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing executable instructions, which, when executed by a computer, enable the method described in the first aspect or any possible implementation thereof.

[0016] One or more technical solutions provided in the embodiments of this application have at least the following technical effects: This application provides a scene target detection method based on a visual saliency model. By clearly defining the detection target and setting constraints, the model focuses more on key targets in the scene during training. This helps improve the model's recognition accuracy of key targets, reduces false positives and false negatives, and thus improves the overall detection accuracy. Training with a dataset containing rich image data and its precise location and category annotations allows the model to learn more diverse feature representations, maintaining high detection performance in different scenes and complex environments. Introducing residual blocks between the convolutional neural network and the multilayer perceptron effectively alleviates the gradient vanishing or exploding problem in deep network training, accelerating the model's training speed and enabling it to process image data faster during inference, thus improving detection efficiency. By distinguishing between salient and non-salient regions in the image through the output layer of the multilayer perceptron, the model can automatically learn and identify the most representative regions in the image, providing more accurate prior information for subsequent target detection and further optimizing the target detection effect. The constructed model framework has good flexibility and scalability. The number of residual blocks, the configuration of convolutional layers, and the structure of the multilayer perceptron can be adjusted according to actual needs to adapt to target detection tasks of varying complexity and scale. This addresses the problem that traditional algorithms often suffer from high false detection rates and poor robustness in complex and ever-changing simulated training environments. In particular, their detection performance deteriorates significantly under conditions such as target occlusion, high background-target similarity, and large variations in lighting conditions, resulting in low saliency values ​​for traditional saliency detection algorithms in complex images and their inability to adapt to different environments and scene changes. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart of a scene target detection method based on a visual saliency model provided in an embodiment of this application; Figure 2 The following is a flowchart illustrating the process of evaluating and optimizing the performance of the initial saliency model on the validation set to obtain the final saliency model, as provided in this embodiment of the application. Figure 3 A detailed flowchart of the multi-scale feature fusion step introduced into the residual block is provided for the embodiments of this application; Figure 4A schematic diagram of a scene target detection device based on a visual saliency model provided in an embodiment of this application; Figure 5 This is a schematic diagram of a scene target detection server based on a visual saliency model provided in an embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0020] The following description of some technologies involved in the embodiments of this application is provided to aid understanding and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, some descriptions of well-known functions and structures are omitted in the following description.

[0021] This application provides a scene target detection method based on a visual saliency model, such as... Figure 1 As shown, the method includes steps S101 to S106. Wherein, Figure 1 This is merely one execution order shown in the embodiments of this application and does not represent the only execution order for a scene object detection method based on a visual saliency model. Where the final result can be achieved, Figure 1 The steps shown can be performed in parallel or in reverse order.

[0022] S101: Collect images from the scene, label the location and category of the targets in the images, and build a dataset.

[0023] Specifically, the scenarios in this application can be very broad. This application takes the scenario of a simulated training battlefield as an example for illustration. Its core value lies in providing a highly simulated, safe and controllable environment to achieve specific training or testing objectives.

[0024] First, images from simulated training battlefield scenarios need to be collected from multiple sources. These images should cover different environmental conditions (e.g., daytime, nighttime, different weather conditions), different target types (e.g., fighter groups, fleets, armored vehicles, airfields, etc.), and different detection difficulties (e.g., target occlusion, complex backgrounds, etc.). When collecting images from the scenarios, it is crucial to ensure data diversity, as this helps the model learn the feature representations of different targets under different environments. Simultaneously, data quality is also critical; high-quality samples can more accurately reflect the true characteristics of the targets, thereby improving the model's evaluation performance.

[0025] Before using collected data for model training, data preprocessing can be performed. The purpose of preprocessing is to improve data quality and enhance the data's feature representation ability, thereby improving the model's training efficiency and detection accuracy. Specifically, preprocessing can include: 1. Image enhancement. By adjusting parameters such as brightness, contrast, and saturation of images, or adding effects such as noise and blurring, image changes under different environmental conditions can be simulated, enhancing the model's adaptability to complex environments. 2. Data transformation. This includes operations such as image scaling, cropping, and rotation to generate more training samples, increase data diversity, and help the model learn more robust feature representations. 3. Data denoising. Removing noise and interference factors from images, such as sensor noise and transmission errors, to improve image quality and enable the model to extract target features more accurately.

[0026] Furthermore, the collected images are annotated in detail, including determining the location (usually using bounding boxes) and category (e.g., fighter jets, armored vehicles) of each salient object in the image. The location of salient objects can be annotated using bounding boxes, and the accuracy of the annotation is crucial for subsequent model training. The annotated images are then organized into a structured dataset. This dataset will be used to train the model and serve as the basis for subsequent model evaluation and optimization.

[0027] S102: Define the detection target and set constraints. The defined detection target is to ensure the accuracy of detecting key targets in the scene, and the set constraints are the salience of the detected targets and the effectiveness of the target detection.

[0028] Specifically, the detection objective is defined as ensuring the accurate detection of key targets in the scene. Key targets are salient objects, such as fighter groups, fleets, and armored vehicles. Constraints are set to require the model to accurately identify salient regions in the image, i.e., targets with salient features that are easily noticed. This requires the model to learn these salient features during training and accurately distinguish them in practical applications. Besides ensuring salientity, constraints also require the model to achieve certain performance standards during detection, including high precision (the proportion of truly salient targets among detected targets), high recall (the proportion of salient targets detected), and a reasonable F1 score (the harmonic mean of precision and recall). Furthermore, the model's classification performance can be evaluated using ROC curves and AUC values.

[0029] S103: Based on the detection target and constraints, construct residual blocks between the convolutional neural network and the multilayer perceptron to introduce residual connections and build an initial saliency model.

[0030] Specifically, when combining Convolutional Neural Networks (CNNs) with Multilayer Perceptrons (MLPs), the spatial hierarchical features that CNNs excel at extracting are partially lost due to the fully connected nature of the MLP. This loss of information affects the model's accurate identification of salient regions in an image, resulting in lower-than-expected image saliency values ​​in the detection results. To overcome this challenge and fully leverage the respective advantages of CNNs and MLPs, this application employs a residual connection strategy. Residual connections not only solve the common gradient vanishing and gradient exploding problems in neural networks but also facilitate the effective transfer and learning of model features.

[0031] Specifically, based on the detection target and constraints, residual blocks are constructed between the convolutional neural network and the multilayer perceptron to introduce residual connections. This includes defining a main path function and a shortcut function. The output of the main path function is added to the output of the shortcut function to obtain the output data of the residual block, thus introducing residual connections.

[0032] Furthermore, the main path is the part of the residual block responsible for feature extraction and transformation. The main path function receives the input data and outputs a feature map after processing it through multiple operations. These operations include convolution, batch normalization, and non-linear activation function operations.

[0033] The main path operation is described as follows: The input data first passes through a convolutional layer, which uses a set of learnable convolutional kernels (or filters) to perform convolution operations on the input data to extract local features. The output of the convolutional layer is a feature map, which preserves the spatial hierarchical information in the input data but abstracts and enhances the features. The output of the convolutional layer is then fed into a batch normalization layer. Batch normalization normalizes each channel of the feature map to have zero mean and unit variance. This helps accelerate the training process, improve the model's convergence speed, and reduce the model's sensitivity to initial weights and learning rate. The output of the batch normalization layer is the normalized feature map. Finally, the normalized feature map passes through a non-linear activation function, typically ReLU (Rectified Linear Unit). The ReLU function introduces non-linearity into the model by setting all negative values ​​to zero and keeping positive values ​​unchanged. This enhances the model's expressive power, enabling it to learn more complex feature maps.

[0034] Therefore, the expression for the main path function is: .in, For input data, Indicates input data Perform convolution operations. This indicates that batch normalization is performed on the output after the convolution operation. This indicates the application of batch normalized output. A nonlinear activation function is used to obtain the feature map. That is, the output of the main path function.

[0035] Specifically, the input data for the shortcut function is the same as that for the main path function. A weight matrix is ​​introduced, and the input data is multiplied by the weight matrix to obtain the output of the shortcut function. (Weight matrix) With input data Dimensional compatibility, used for input data Perform a linear transformation. This transformation is achieved through matrix multiplication. The expression for the shortcut function is: .in, The input data is the same as the input data of the main path function. This is the weight matrix. This is the output of the shortcut function.

[0036] weight matrix It is a learnable parameter matrix that is optimized during training to minimize the network's overall loss. This is achieved by adjusting the weight matrix. The value can help learn how to best process input data. Convert into a shortcut function output representation The output of this shortcut function is then added to the output of the main path function to obtain the output data of the residual block. Weight matrix Initially, the data is initialized with small random numbers. These random numbers are typically random and close to zero to avoid gradient explosion or vanishing during the early stages of training. During training, the input data... First, a series of complex transformations are performed using the main path function, while the shortcut function also calculates its output. Then, the outputs of the main path function and the shortcut function are added together to obtain the output data of the residual block. This output data continues to propagate forward until the network produces a prediction. The network's prediction is compared with the true labels, and the loss (e.g., cross-entropy loss) is calculated. The loss value is then backpropagated back through each layer of the network using the backpropagation algorithm. The gradient of the loss value with respect to the shortcut function output is further propagated to the weight matrix. Specifically, the loss with respect to the weight matrix will be calculated according to the chain rule. The gradient, i.e. ,in, The loss function L is represented by the weight matrix. The partial derivatives are then calculated. Finally, optimization algorithms (such as SGD, Adam, etc.) are used to update the weight matrix based on the gradient information. The value of . The update process aims to reduce the loss and make the network's predictions closer to the true labels.

[0037] S104: Obtain the output data of the residual block as the input data of the multilayer perceptron.

[0038] Specifically, since the output data of the residual block is the sum of the output of the main path function and the output of the shortcut function, the expression for the output data of the residual block is: .in, This is the output data for the residual block. The output of the shortcut function, This is the output of the main path function.

[0039] It's important to note that if the dimension of the residual block's output data doesn't match the expected input dimension of the multilayer perceptron, a dimensionality transformation (such as flattening, reshaping, or additional convolutional layers) is required. Once the dimensions match, the residual block's output data can be used as input to the multilayer perceptron. The first layer (and subsequent layers) of the multilayer perceptron will further process this input to extract higher-level features. The expression for the multilayer perceptron is: .in, This is the output data for the residual block. This is the weight matrix of the multilayer perceptron. The weight matrix takes the output of the residual block as input and performs a linear transformation on it. The bias term is an additional parameter for each neuron (or node) in a multilayer perceptron that allows the model to have a non-zero output even without any input.

[0040] S105: Output the first prediction result of the initial saliency model through the output layer of the multilayer perceptron. The first prediction result is used to distinguish between salient and non-salient regions of the image.

[0041] Specifically, the primary objective of the first prediction result is to distinguish between salient and non-salient regions in an image. A salient region refers to a prominent part of the image, which can be the target object of interest or background information closely related to the target object. Through the output layer of the multilayer perceptron, a saliency map of the same size as the input data can be output. In this saliency map, the value of each pixel represents the probability that the pixel belongs to a salient region. By setting a threshold, the saliency can be further classified. Figure 2 The saliency is converted into a mask to distinguish between salient and non-salient regions in the image. Therefore, the first prediction result refers to the saliency map or saliency mask obtained after the multilayer perceptron performs preliminary processing and analysis on the input image. This prediction result is a preliminary judgment of the image's saliency and can be used in subsequent model evaluation and optimization.

[0042] S106: Train the initial saliency model using the dataset, evaluate and optimize the performance of the initial saliency model on the validation set, and thus obtain the final saliency model.

[0043] Figure 2 The flowchart provided in this application embodiment describes the evaluation and optimization of the performance of the initial saliency model on the validation set to obtain the final saliency model, as follows: Figure 2 As shown, it includes steps S201 to S203.

[0044] S201: In the experimental tools for equipping intelligent algorithms, design evaluation metrics and construct representative scenarios as simulation scenarios. The evaluation metrics include detection accuracy, false alarm rate, false negative rate, and processing speed. Representative scenarios include those that can cover different environmental conditions, different target types, and different detection difficulties.

[0045] Specifically, a comprehensive set of evaluation metrics needs to be designed in the equipment intelligent algorithm experimental tool to fully measure the performance of the saliency model. The equipment intelligent algorithm experimental tool can be TensorFlow (an open-source machine learning library). Further, the evaluation metrics include detection accuracy (the ratio of correctly detected targets to the total number of targets), false positive rate (the proportion of non-targets incorrectly identified as targets), false negative rate (the ratio of undetected targets to the total number of targets), and processing speed (the speed at which the model processes the input image and provides detection results). Simultaneously, a set of representative simulation scenarios needs to be constructed. These scenarios should cover different environmental conditions (such as changes in lighting, weather conditions, etc.), different target types (such as vehicles, personnel, buildings, etc.), and different detection difficulties (such as target occlusion, small target detection, etc.). Constructing simulation scenarios also includes editing the scenarios, generating the equipment system participating in the experiment, and setting up data interaction interfaces. By constructing simulation scenarios, the comprehensiveness and reliability of the evaluation results can be ensured. Setting up data interaction interfaces can accurately respond to changes in the experimental scenarios, i.e., the simulation scenarios, and transmit data in real time.

[0046] S202: Load the simulation scenario and initial saliency model in the virtual independent operating environment, conduct simulation experiments, and collect and store the experimental data in real time.

[0047] Specifically, in a virtual independent runtime environment, the pre-built simulation scenario and initial saliency model are loaded. The simulation experiment is initiated by pushing the simulation scenario data and trigger commands for the saliency model algorithm to the virtual independent runtime environment, allowing the initial saliency model to detect targets in the simulation scenario within the virtual environment. During the experiment, experimental data can be collected and stored in real time through a data center. This data includes the detection results of the initial saliency model, processing time, and relevant scenario information.

[0048] Furthermore, it is also possible to load significance models that need to be evaluated, including significance models with residual connections (i.e., the initial significance model) and significance models without residual connections.

[0049] S203: After the experiment, the collected experimental data are evaluated and analyzed, and the initial significance model is adjusted and optimized based on the evaluation and analysis results to obtain the final significance model.

[0050] Specifically, after the experiment, the collected experimental data is used for evaluation and analysis. Based on the evaluation and analysis results (such as detection accuracy, false alarm rate, and false negative rate), the problems and shortcomings of the initial significance model can be identified. Subsequently, the model is adjusted and optimized to address these problems, which may include modifying the objective function, adjusting constraints, and introducing new variables or parameters. The adjusted model will then be re-simulated to verify the optimization effect. This process will be iterative until a final significance model that meets the performance requirements is obtained.

[0051] The scene target detection method based on the visual saliency model provided in this application embodiment further includes: introducing a multi-scale feature fusion step in the residual block.

[0052] Figure 3 A detailed flowchart illustrating the multi-scale feature fusion step introduced into the residual block is provided for embodiments of this application, as follows: Figure 3 As shown, it includes steps S301 to S308.

[0053] S301: In the main path function, multiple convolutional kernels of different scales are used to perform convolution operations on the input data to generate feature maps of multiple scales.

[0054] Specifically, in the direct path of the main path function, i.e., the residual block, multiple convolutional kernels with different scales can be used to convolve the input data. These kernels can be 3x3, 5x5, etc., to capture local and global information at different scales. This step generates feature maps at multiple scales, each containing feature representations of the input data at different scales.

[0055] Furthermore, the input data in this application is image data, defined as `input`, which is typically a three-dimensional array (or tensor) with dimensions H (height), W (width), and C (number of channels). The input data can be a grayscale image (C=1) or a color image (C=3). A set of multi-scale convolutional kernels is defined, such as kernel1, kernel2, ..., kernelN, where N is the total number of kernels. Each kernel `kerneli` (i from 1 to N) has its own unique size (e.g., 3*3, 5*5, 7*7, etc.) and depth (usually the same as the number of channels C of the input image). These kernels have different sizes, enabling them to capture features at different scales in the input image. For each kernel `kerneli`, a convolution operation `conv` is performed to generate the corresponding feature map. Therefore, the expression for the feature map is: `Fi = conv(input, kernelli)`. The convolution operation can be specifically described as follows: the convolution kernel slides across the input image. For each position, the sum of the element-wise products of the kernel and the corresponding region (also called the receptive field) of the input image is calculated, and then an optional bias term is added. This process generates a new two-dimensional array, the feature map, whose size depends on the kernel size, stride, and padding settings. Since the goal of this application is to ensure accurate detection of key targets in a scene (such as fighter groups, fleets, armored vehicles, etc.), the kernel size setting should take into account the possible scale and complexity of these targets in the image. When using a 3x3 kernel, its small size allows it to capture local details in the input image, such as edges and textures. When using a 5x5 kernel, it has a larger receptive field, thus capturing information from a larger area. Since key targets may have different scales in the image, the kernel size setting should be able to cover these different scales. This can be achieved by setting multiple kernels of different sizes or using convolutional layers with different receptive fields. Through convolution, a set of feature maps is obtained, each of which contains feature representations of the input image at different scales.

[0056] S302: Perform non-uniform segmentation on the generated feature map at each scale to divide it into multiple regions of interest of different scales and shapes, so that each feature map has multiple regions of interest.

[0057] Specifically, the goal of non-uniform segmentation is to divide a feature map into multiple regions, each of which is a Region of Interest (ROI). "Non-uniform" here means that the size, shape, and location of the RIO can be irregular on the feature map to better accommodate targets or features of different scales and shapes in the image. For example, for a fighter jet group, multiple small RIOs can be generated to capture each fighter jet individually; while for a fleet or armored vehicles, larger and irregularly shaped RIOs can be generated to cover the entire target area. RIOs can have different sizes and shapes to accommodate various targets or features that may appear in the image. For example, some RIOs may be rectangular to capture regular targets, while others may be irregularly shaped to capture objects with complex boundaries. RIOs can overlap, so that even if a target or feature spans multiple RIOs, it can still be adequately captured and represented. Overlapping RIOs can provide multiple perspectives and feature representations of the same region, helping to enhance the robustness and accuracy of the model. For example, even if the target's position, pose, or scale changes in the image, the RIO can still capture the target's features. Specific segmentation strategies can be designed according to task requirements and the properties of the feature map. For example, a sliding window method combined with windows of different scales can be used to generate candidate regions of interest; or image segmentation algorithms (such as graph-based segmentation, level set methods, etc.) can be used to directly generate irregularly shaped regions of interest. The expression for non-uniform segmentation is: .in, This refers to the region of interest formed after a non-uniform segmentation operation. To make feature maps The operation of dividing into multiple regions of interest.

[0058] S303: Determine the relative position of each region of interest on the same scale feature map and map it onto its respective feature map.

[0059] S304: Using the center point of each region of interest on the feature map as the center, extract a small region around the center point as the mapped feature map.

[0060] Specifically, the detection objective defined in this application is to ensure the accurate detection of key targets in the scene, such as building complexes. These targets may have different scales and shapes and need to be detected at different feature scales. Therefore, it is first necessary to determine the feature map of each region of interest. The relative position of the region of interest (ROI) on the feature map. This typically involves calculating the coordinates of reference points such as the bounding box or center point of the ROI in the feature map coordinate system. These coordinates can be pixel-level or values ​​that have undergone some normalization or scaling. Once the relative position of the ROI on the feature map is known, a mapping operation can be performed. The purpose of mapping is to associate the ROI with a corresponding region on the feature map so that features can be extracted from that region. If the coordinates of the center point of the ROI on the feature map are (x, y), a rectangular region centered at (x, y) can be selected as the mapped feature map region. The size of this rectangular region can be determined based on the actual size of the ROI and the resolution of the feature map to ensure that the ROI is adequately covered. Another method is to directly use the bounding box coordinates of the ROI for mapping, mapping the coordinates of the top-left and bottom-right corners (or two other diagonal points) of the ROI onto the feature map, and determining a rectangular region accordingly as the mapped feature map region. The expression for the mapping operation is: .in, It is a region of interest formed after a non-uniform segmentation operation. Mapping to feature map Then generate the mapped feature map. The mapped feature map can be used to determine the presence of a target in the region of interest and to estimate the target's category and location. By mapping regions of interest onto feature maps of the same scale, specific image regions can be processed and analyzed more effectively.

[0061] S305: Perform a convolution operation on each mapped feature map to generate a new feature map with the same number of regions of interest.

[0062] Specifically, after mapping the region of interest onto a feature map, a convolution operation can be performed on each mapped feature map. The convolution operation can further extract useful information from the mapped feature maps. The expression for the convolution operation is: .in, For convolution operations, This is the convolution kernel used to extract important information from the mapped feature map. The mapped feature map, This is the new feature map obtained after the convolution operation. It's important to note that the convolution operation here is performed on each mapped feature map. Since these steps are performed independently, a new set of feature maps will eventually be obtained that is the same number as the number of regions of interest.

[0063] S306: Perform a fusion operation on multiple new feature maps to generate the final feature map.

[0064] Specifically, after extracting useful information from each mapped feature map through convolution operations, a fusion operation is needed to synthesize this information to form a more comprehensive and accurate description of the region of interest. The expression for the fusion operation is: .in, For the new feature map, For the final feature map, This indicates a fusion operation on the new feature map, which can be implemented using various strategies. The final feature map obtained through the fusion operation integrates feature information from multiple regions of interest.

[0065] S307: Use the final feature map as input data for the multilayer perceptron.

[0066] Specifically, after fusing multiple new feature maps, a final feature map integrating information from multiple regions of interest is obtained. This final feature map contains rich visual features. To extract higher-level semantic information from these feature maps and transform it into prediction results that can be used for object detection tasks, this final feature map is used as input data for a multilayer perceptron.

[0067] A multilayer perceptron is a feedforward neural network consisting of multiple layers, each containing multiple neurons. These neurons are interconnected through weighted connections, and non-linear activation functions can be introduced to increase the network's complexity. By feeding the final feature map into the multilayer perceptron, the network's powerful learning capabilities can be leveraged to capture important information from the feature map and transform it into a more meaningful representation.

[0068] S308: Output the second prediction result of the initial saliency model through the output layer of the multilayer perceptron. The second prediction result is used to distinguish between salient and non-salient regions of the image.

[0069] Specifically, after the multilayer perceptron processes the final feature map, its output layer produces a second prediction result from the initial saliency model. The main purpose of the second prediction result is to distinguish between salient and non-salient regions of the image. Salient regions typically refer to the most striking parts of the image; they can be the target object of interest or background information closely related to the target object. The second prediction result helps the initial saliency model better understand the image content. Simultaneously, the second prediction result can also serve as an important metric for model performance evaluation; by comparing it with ground truth labeled data, the model's accuracy and robustness can be assessed.

[0070] This application also provides a scene target detection device 400 based on a visual saliency model, such as... Figure 4As shown, the device includes: a collection module 401, a definition module 402, a construction module 403, an acquisition module 404, an output module 405, and an optimization module 406.

[0071] The collection module 401 is used to collect images in the scene, label the location and category of the targets in the images, and build a dataset.

[0072] The definition module 402 is used to define the detection target and set constraints.

[0073] The building module 403 is used to build residual blocks between the convolutional neural network and the multilayer perceptron according to the detection target and constraints, so as to introduce residual connections and build an initial saliency model.

[0074] The acquisition module 404 is used to acquire the output data of the residual block as the input data of the multilayer perceptron.

[0075] The output module 405 is used to output the first prediction result of the initial saliency model through the output layer of the multilayer perceptron. The first prediction result is used to distinguish between salient and non-salient regions of the image.

[0076] The optimization module 406 is used to train the initial saliency model using the dataset, evaluate and optimize the performance of the initial saliency model on the validation set, and thus obtain the final saliency model.

[0077] Some modules in the apparatus described in this application can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0078] The apparatus or module described in the above embodiments can be implemented by a computer chip or physical entity, or by a product with a certain function. For ease of description, the above apparatus is described by dividing it into various modules according to their functions. When implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.

[0079] The methods, apparatus, or modules described in this application can be implemented in a computer-readable program code manner. The controller can be implemented in any suitable manner, for example, as a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of a memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code manner, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included within it for implementing various functions can also be considered as structures within the hardware component. Alternatively, the device used to implement various functions can be viewed as either a software module that implements the method or a structure within a hardware component.

[0080] like Figure 5 As shown in the figure, this application embodiment also provides a scene target detection server based on a visual saliency model, including a memory 501 and a processor 502; the memory 501 is used to store computer-executable instructions; the processor 502 is used to execute the computer-executable instructions to implement the scene target detection method based on a visual saliency model described above in this application embodiment.

[0081] This application also provides a computer-readable storage medium storing executable instructions, which, when executed by a computer, enable the scene target detection method based on a visual saliency model described above in this application.

[0082] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, or it can be embodied in the process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in the embodiments of this application.

[0083] The various embodiments described in this specification are presented in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. All or part of this application can be used in numerous general-purpose or special-purpose computer system environments or configurations.

[0084] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of this application.

Claims

1. A scene target detection method based on a visual saliency model, characterized in that, include: Collect images from the scene, label the location and category of the targets in the images, and build a dataset; Define the detection target and set constraints; Based on the detection target and constraints, residual blocks are constructed between the convolutional neural network and the multilayer perceptron to introduce residual connections and build an initial saliency model; The output data of the residual block is obtained as the input data of the multilayer perceptron; The first prediction result of the initial saliency model is output through the output layer of the multilayer perceptron; wherein, the first prediction result is used to distinguish between salient and non-salient regions of the image; The initial saliency model is trained using the dataset, and its performance is evaluated and optimized on the validation set to obtain the final saliency model. The step of constructing residual blocks between the convolutional neural network and the multilayer perceptron to introduce residual connections, based on the detection target and constraints, includes: Define the main path function and the shortcut function; The output of the main path function is added to the output of the shortcut function to obtain the output data of the residual block, which is then used to introduce residual connections.

2. The scene target detection method based on the visual saliency model according to claim 1, characterized in that, The main path function is used to receive input data and output a feature map after performing multiple operations on the input data. Multiple operations include convolution, batch normalization, and non-linear activation function operations.

3. The scene target detection method based on the visual saliency model according to claim 2, characterized in that, The input data of the shortcut function is the same as the input data of the main path function; By introducing a weight matrix and multiplying the input data with the weight matrix, the output of the shortcut function can be obtained.

4. The scene target detection method based on the visual saliency model according to claim 1, characterized in that, Also includes: A multi-scale feature fusion step is introduced into the residual block.

5. The scene target detection method based on the visual saliency model according to claim 4, characterized in that, The multi-scale feature fusion step includes: The main path function uses multiple convolutional kernels of different scales to perform convolution operations on the input data, generating feature maps of multiple scales; The generated feature map at each scale is divided into multiple regions of interest of different scales and shapes by performing a non-uniform segmentation operation, so that each feature map has multiple regions of interest. Determine the relative position of each region of interest on the same scale feature map and map it to its respective feature map; Using the center point of each region of interest on the feature map as the center, a small region is extracted around the center point as the mapped feature map. Perform a convolution operation on each mapped feature map to generate a new feature map with the same number of regions of interest. Multiple new feature maps are fused to generate the final feature map. The final feature map is used as input data for the multilayer perceptron. The second prediction result of the initial saliency model is output through the output layer of the multilayer perceptron; wherein the second prediction result is used to distinguish between salient and non-salient regions of the image.

6. The scene target detection method based on the visual saliency model according to claim 1, characterized in that, The step of evaluating and optimizing the performance of the initial saliency model on the validation set to obtain the final saliency model includes: In the experimental tool for equipping intelligent algorithms, evaluation indicators are designed and representative scenarios are constructed as simulation scenarios; wherein, the evaluation indicators include detection accuracy, false alarm rate, false negative rate and processing speed, and the representative scenarios include scenarios that can cover different environmental conditions, different target types and different detection difficulties; The simulation scenario and initial saliency model are loaded into a virtual independent operating environment to conduct simulation experiments and collect and store experimental data in real time. After the experiment, the collected experimental data were evaluated and analyzed, and the initial significance model was adjusted and optimized based on the evaluation and analysis results to obtain the final significance model.

7. A scene target detection device based on a visual saliency model, characterized in that, include: The collection module is used to collect images from the scene, label the location and category of targets in the images, and build a dataset; The definition module is used to define the detection target and set constraints. The building module is used to construct residual blocks between the convolutional neural network and the multilayer perceptron based on the detection target and constraints, so as to introduce residual connections and build an initial saliency model; An acquisition module is used to acquire the output data of the residual block as input data for the multilayer perceptron; The output module is used to output the first prediction result of the initial saliency model through the output layer of the multilayer perceptron; wherein the first prediction result is used to distinguish between salient and non-salient regions of the image; An optimization module is used to train the initial saliency model using the dataset, evaluate and optimize the performance of the initial saliency model on the validation set, and thus obtain the final saliency model. The step of constructing residual blocks between the convolutional neural network and the multilayer perceptron to introduce residual connections, based on the detection target and constraints, includes: Define the main path function and the shortcut function; The output of the main path function is added to the output of the shortcut function to obtain the output data of the residual block, which is then used to introduce residual connections.

8. A scene target detection server based on a visual saliency model, characterized in that, Including memory and processor; The memory is used to store computer-executable instructions; The processor is configured to execute the computer-executable instructions to implement the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores executable instructions, which, when executed by a computer, enable the implementation of the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Panoramic image target detection method and system based on improved YOLOv7

    CN116665007A

  • Salient target detection method based on light field refocusing data enhancement

    CN116778187A