Data processing method and device applied to equipment detection, equipment and program product

By building a target detection model, combining image and text encoders to extract features and perform bidirectional alignment encoding, the detection confusion problem caused by the similar appearance of network devices is solved, and the accuracy of device detection is improved.

CN120670930APending Publication Date: 2025-09-19CHINA MOBILE GROUP ZHEJIANG +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510621183.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the existing technology of network equipment quality inspection, due to the high similarity in the appearance design of equipment of different brands or models, it is difficult for the target detection algorithm to extract effective discriminative features, resulting in false detection or missed detection, affecting the accuracy and reliability of quality inspection.

Method used

Build an object detection model, including an image encoder, a text encoder, a feature fusion device, and a detector. The image encoder extracts visual features, the text encoder extracts semantic description features, and the feature fusion device performs bidirectional alignment encoding to generate enhanced semantic description features and image features. The detector is trained with labels to generate detection results.

Benefits of technology

It effectively alleviates the problem of confusion between visually similar features, significantly improves the accuracy of device detection, and can focus on the semantic differences between devices, thereby improving the accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670930A_ABST
    Figure CN120670930A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device applied to equipment detection, equipment and a program product. According to the scheme, a target detection model comprising an image encoder, a text encoder, a feature fusion device and a detector is specifically constructed. In the training process of the target detection model, firstly, a sample image of target equipment is encoded through an image encoder, and image features of the target equipment are extracted; meanwhile, semantic coding is conducted on the descriptive text of the target device through a text coder, and semantic description features are generated; then, the feature fusion device carries out bidirectional alignment coding on the image features and the semantic description features to generate enhanced image features and semantic description features; the enhanced features are used as input conditions of a detector, and the detector is trained in combination with tags (target device types and detection frame positions) corresponding to sample images, so that semantic information features and visual features are deeply fused to provide accuracy of device-by-device detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of equipment detection, and in particular to a data processing method, apparatus, equipment and program product applied to equipment detection. Background Art

[0002] With the rapid deployment and intelligent transformation of 5G networks, automated quality inspection of production equipment has become a core requirement for efficient operation and maintenance of communications infrastructure. Currently, mainstream solutions mainly use deep learning-based object detection algorithms (such as YOLOv5 and Faster R-CNN) to achieve device positioning, but they face the following challenges in practical applications: Because network devices of different brands or models are highly similar in appearance design (such as slight differences in visual features such as shape, interface layout, and color), the differences within the class are even smaller than the differences between classes, making it difficult for the algorithm to extract effective discriminative features, which in turn leads to problems such as false detection (identifying non-target devices as targets) or missed detection (ignoring real targets), seriously restricting the accuracy and reliability of quality inspection. Summary of the Invention

[0003] In view of the above problems, the present application provides a data processing method, apparatus, device, and program product for device detection that overcomes or at least partially solves the above problems. The technical solution is as follows: In a first aspect, a data processing method for device detection is provided, comprising: Building an object detection model, wherein the object detection model includes an image encoder, a text encoder, a feature fuser, and a detector; Encoding a sample image of a target device using the image encoder to obtain image features of the target device; and semantically encoding a device description text of the target device using the text encoder to obtain semantic description features of the target device; Performing bidirectional alignment encoding on the image features and the semantic description features by the feature fusion device to obtain enhanced image features and enhanced semantic description features; The enhanced semantic description feature is used as the conditional input of the detector, and the detector is trained based on the enhanced image feature and the label corresponding to the sample image; wherein the enhanced semantic description feature is used to generate the initial learnable parameters of the detector, and the detector generates a detection result by iterating the learnable parameters, and the detection result includes the detection box position and the device type; the label is marked with the device type of the target device and the detection box position corresponding to the target device in the sample image.

[0004] In a second aspect, a data processing apparatus for device detection is provided, comprising: A model building module is used to build an object detection model, wherein the object detection model includes an image encoder, a text encoder, a feature fuser and a detector; a feature encoding module, configured to encode a sample image of a target device using the image encoder to obtain image features of the target device; and to semantically encode a device description text of the target device based on the text encoder to obtain semantic description features of the target device; A feature fusion module, configured to perform bidirectional alignment encoding on the image features and the semantic description features through the feature fusion device to obtain enhanced image features and enhanced semantic description features; A model training module is used to use the enhanced semantic description feature as the conditional input of the detector and train the detector based on the enhanced image feature and the label corresponding to the sample image; wherein the enhanced semantic description feature is used to generate the initial learnable parameters of the detector, and the detector generates a detection result by iterating the learnable parameters, and the detection result includes the detection box position and the device type; the label is marked with the device type of the target device and the detection box position corresponding to the target device in the sample image.

[0005] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor; and a memory configured to store computer-executable instructions, wherein the computer-executable instructions, when executed, cause the processor to execute the method described in the first aspect.

[0006] According to a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium is used to store computer-executable instructions, and the computer-executable instructions implement the method described in the first aspect when executed by a processor.

[0007] The present embodiment constructs an object detection model to detect devices. The object detection model includes four core components: an image encoder, a text encoder, a feature fusion unit, and a detector. First, the image encoder encodes a sample image of the target device and extracts image features containing key information such as color, texture, and shape. Simultaneously, the text encoder semantically encodes the device's description text to generate textual semantic description features. These semantic description features can reflect information such as the device's brand and model. Next, the feature fusion unit fuses the image features and semantic description features through bidirectional alignment to generate enhanced semantic description features and enhanced image features. During the model training phase, the detector uses the enhanced semantic description features as input and combines them with the labels corresponding to the sample images. By iteratively optimizing the learning parameters, it generates detection results that include the target box location and device type. This reinforcement mechanism based on textual semantic supervision enables the object detection model to focus on semantic differences between devices, such as different descriptions of brands or types, effectively alleviating the problem of feature confusion between visually similar features and significantly improving the accuracy of device detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0009] Figure 1 This is a flow chart of a data processing method applied to device detection according to an embodiment of the present application.

[0010] Figure 2 This is a first structural diagram of the target detection model of the data processing method of an embodiment of the present application.

[0011] Figure 3 This is a second structural diagram of the target detection model of the data processing method of an embodiment of the present application.

[0012] Figure 4 This is a third structural diagram of the target detection model of the data processing method of an embodiment of the present application.

[0013] Figure 5 This is a structural diagram of a data processing device applied to equipment detection according to an embodiment of the present application.

[0014] Figure 6 This is a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0015] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this specification.

[0016] This application proposes a data processing solution for device detection, which aims to solve the detection confusion problem caused by high target similarity in network device quality inspection.

[0017] Specifically, the data processing solution of the present application involves data processing, apparatus, equipment, and program products applied to device detection. The following describes each embodiment in detail.

[0018] An embodiment of the present application provides a data processing method for device detection. Figure 1 Schematic diagram of the data processing method, including the following steps: S102, constructing a target detection model, which includes an image encoder, a text encoder, a feature fuser, and a detector.

[0019] The target detection model structure of this embodiment is as follows Figure 2 Unlike traditional object detection models, which only accept image data as input and configure an image encoder to extract image features, the object detection model of this embodiment not only takes image data as input but also introduces text input and configures a text encoder.

[0020] S104 , encoding the sample image of the target device through an image encoder to obtain image features of the target device; and semantically encoding the device description text of the target device based on a text encoder to obtain semantic description features of the target device.

[0021] This embodiment uses an image encoder to encode a sample image of the target device and extract image features such as color, texture, and shape. Simultaneously, a text encoder performs semantic encoding on the device description text of the target device and extracts semantic description features such as device type information and brand information. The specific implementation is as follows: 1) Image Encoder Input: Sample image of target device (size: ).

[0022] Encoder architecture: Swin-Transformer is used as the image encoder, and multi-scale image features (atlas) are generated through hierarchical feature extraction: Output scale: 、 、 、 .

[0023] Function: Capture key visual features of target devices (such as interface layout, shape, and color) in sample images.

[0024] 2) Text Encoder Input: Device description text, including device type and brand information.

[0025] Encoder architecture: A pre-trained language model (BERT or LSTM) is used to extract semantic embedding vectors, which are then combined with a text feature extraction module to generate structured semantic vectors, i.e., semantic description features.

[0026] Output format: semantic description feature dimension is ,in, is the length of the text sequence, is the feature dimension.

[0027] Function: Implant semantic description features to distinguish them from image features, thereby preventing feature confusion caused by high visual similarity.

[0028] S106 , performing bidirectional alignment encoding on the image features and the semantic description features through a feature fusion device to obtain enhanced image features and enhanced semantic description features.

[0029] This embodiment is performed through a feature fuser: encoding image features based on a first cross-attention mechanism to obtain enhanced image features; encoding semantic description features based on a second cross-attention mechanism to obtain enhanced semantic description features; wherein the key vector and value vector of the first cross-attention mechanism are determined based on the semantic description features, and the query vector is determined based on the image features; the key vector and value vector of the second cross-attention mechanism are determined based on the enhanced image features, and the query vector is determined based on the semantic description features.

[0030] The above-mentioned attention mechanism is an important technology in deep learning. Its core idea is to calculate the correlation between different parts of the input data and assign weights to each part, thereby highlighting the most important information for the current task.

[0031] Among them, the attention mechanism involves three important parameters: Q (Query): Query vector, used to represent the input information that needs attention.

[0032] K (Key): Key vector, used to represent the features in the input information.

[0033] V (Value): Value vector, representing the specific content of the input data.

[0034] Before encoding the image features and semantic description features respectively through the feature fusion device, this embodiment can be performed through the feature fusion device: encoding the image features based on the deformable attention mechanism to obtain a first encoding result; encoding the semantic description features based on the self-attention mechanism to obtain a second encoding result; wherein, the key vector and value vector of the first cross-attention mechanism are determined based on the second encoding result, and the query vector is determined based on the first encoding result; the key vector and value vector of the second cross-attention mechanism are determined based on the first encoding result, and the query vector is determined based on the second encoding result.

[0035] As an example introduction, Figure 3 : This is a network structure diagram of the feature fusion device of this embodiment. The corresponding process of generating enhanced image features and enhanced semantic description features includes: 1) Feature preprocessing Image feature adjustment: Multi-scale image features (size ) The number of channels is unified by 1×1 convolution , and expands to The sequence form of .

[0036] Semantic feature alignment: maintaining the semantic feature dimension of the text , spatially aligned with image features.

[0037] 2) Encoding image features based on deformable self-attention mechanism and encoding semantic description features based on self-attention mechanism ①Calculation formula of deformable self-attention mechanism:

[0038] parameter: is the number of attention heads; The number of neighboring features selected in each attention head; is the query vector; is the position code, For the attention head, query location, The offset corresponding to each adjacent feature; For the attention head, query location, The attention weight of the neighboring features corresponding to the current position; and All are The weight matrix corresponding to each attention point.

[0039] It should be noted that if there are multiple sample images with corresponding sizes, the calculation formula of the deformable self-attention mechanism can be converted to:

[0040] parameter: is the size number corresponding to the sample image; For the attention head, Size, query location, The offset corresponding to each adjacent feature; For the attention head, Size, query location, The attention weight of the neighboring features corresponding to the current position; No. The normalized atlas position of the image features of each size after normalization.

[0041] ②Calculation formula of self-attention mechanism:

[0042] Parameters: Q is the query vector; K is the key vector; V is the value vector.

[0043] 3) Dual-path cross attention mechanism ① First cross attention mechanism (picture → text):

[0044] Input: semantic description features As query vector, image features As key vector and value vector.

[0045] Function: Align image features to semantic space and generate semantically enhanced image features.

[0046] ② Second Cross Attention Mechanism (Text → Picture):

[0047] Input: Image features As the query vector, the semantic description feature t is used as the key vector and the value vector.

[0048] Function: Align semantic features to visual space and generate visually guided semantic description features 4) Feature enhancement and output Enhanced image features: The output of the first cross-attention mechanism is processed by a feed-forward network (FFN) to retain details of key areas of the device (such as interfaces and logos).

[0049] Enhance semantic features: The output of the second cross-attention mechanism is processed by FFN to enhance the semantic information associated with the visual features (such as brand model).

[0050] S108, using the enhanced semantic description feature as a conditional input of the detector, and training the detector based on the enhanced image feature and the label corresponding to the sample image; wherein the enhanced semantic description feature is used to generate the initial learnable parameters of the detector, and the detector generates a detection result by iterating the learnable parameters, and the detection result includes the detection box position and the device type; the label is marked with the device type of the target device and the detection box position corresponding to the target device in the sample image.

[0051] In detector technology, object queries, or the learnable parameters described in this embodiment, are used to represent the model's detection hypotheses about potential targets in an image. During detector training, the learnable parameters, as training objects, are continuously adjusted to better match the actual target location and category. Correspondingly, this embodiment enhances semantic description features to generate the detector's initial learnable parameters. This allows the initialized learnable parameters to be imbued with semantic information related to brand, model, and other information, providing directional guidance at the initial stages of training.

[0052] For example, suppose the detection task is to distinguish between "Brand A optical synchronization network equipment" and "Brand B packet transmission network equipment" in the computer room.

[0053] The traditional detector training method is to first randomly initialize the learnable parameters, and then learn the visual differences between the two types of equipment (such as specific textures or interface layouts) through a large number of samples. However, since "Brand A's optical synchronization network equipment" and "Brand B's packet transmission network equipment" have similar visual features, it is difficult for the detector to focus on distinguishing features, and recognition confusion is still prone to problems after training.

[0054] The detector training method of this embodiment is: the brand information of "Brand A" or "Brand B" already included in the enhanced semantic description features is embedded into the initial learnable parameters, so that the detector can quickly learn the semantic differences, thereby quickly distinguishing "Brand A's optical synchronization network equipment" and "Brand B's packet transmission network equipment", thereby reducing the probability of confusion.

[0055] As an exemplary introduction, this embodiment can first randomly generate a learnable parameter according to the existing detector training method, and then fuse the enhanced semantic description feature with the randomly generated learnable parameter (for example, convert the two into vectors of the same dimension and add them together) to obtain the initial learnable parameter.

[0056] Specifically, refer to Figure 4 As shown, the detector of this embodiment includes: Multiple layers of serially arranged decoder modules are used in each decoder module. Each decoder module performs the following operations: using a cross-attention mechanism to interact the learnable parameters input to the current layer with the enhanced image features to update the learnable parameters; performing a nonlinear transformation on the updated learnable parameters through a feedforward neural network to generate and output the iterated learnable parameters; and between two adjacent decoder modules, the iterated learnable parameters output by the decoder module of the previous layer are directly used as the learnable parameters input to the decoder module of the next layer. Classification network, predicts device type based on the learnable parameters of the final iteration; input: learnable parameters output by the decoder of the last layer. Output: probability distribution of device type, dimension ;in, represents the amount of sample data, Indicates the number of prediction boxes, It is represented as the number of predicted categories + 1 (which may include background categories).

[0057] The regression network predicts the position of the detection box based on the learnable parameters of the final iteration. Input: The learnable parameters output by the decoder of the last layer. Output: The coordinates of the detection box, with dimensions , Indicates the center coordinates, width and height of the prediction box.

[0058] The training process is as follows: 1) Matching degree calculation: According to the ground truth box provided by each label , calculate its matching degree with the predicted box σ(i) output by training:

[0059] in, for exist Predicted values ​​in categories, To predict the box loss, a variety of regression losses can be used to calculate it, including L1 loss, gIoU loss, etc.

[0060] 2) Hungarian matching: The Hungarian algorithm is used to assign a unique predicted box to each true box to ensure one-to-one matching and avoid redundant detection.

[0061] 3) Determine the loss function: The total loss function is composed of the weighted category loss and coordinate loss:

[0062] Category loss: Only the cross entropy loss of the matching prediction box is calculated, and the remaining prediction boxes are marked as background classes.

[0063] Coordinate loss: The weighted sum of L1 loss and GIoU loss is calculated only for the matched prediction boxes.

[0064] 4) Parameter adjustment: In order to reduce the total loss function, the parameters of the target detection model are adjusted. The adjustment objects must at least include the detector, which can include the image encoder, text encoder, and feature fuser.

[0065] It should be understood that after the detector is trained, the object detection model can be put into use. For example, in this embodiment, a surveillance image of a target device can be input into the target detection model along with the surveillance image and device description text to obtain a target detection result. The target detection result includes the target device type and the corresponding detection box position of the target device in the surveillance image.

[0066] In summary, the method of this embodiment constructs an object detection model to detect devices. The object detection model comprises four core components: an image encoder, a text encoder, a feature fusion unit, and a detector. First, the image encoder encodes a sample image of the target device, extracting image features containing key information such as color, texture, and shape. Simultaneously, the text encoder semantically encodes the device's description text to generate textual semantic features. These features reflect information such as the device's brand and model. Next, the feature fusion unit fuses the image features and semantic features through bidirectional alignment to generate enhanced semantic features and enhanced image features. During the model training phase, the detector uses the enhanced semantic features as input and combines them with the corresponding labels of the sample images. Learning parameters are iteratively optimized to generate detection results that include the target box location and device type. This reinforcement mechanism, based on textual semantic supervision, enables the object detection model to focus on semantic differences between devices, such as different descriptions of brand or type, effectively alleviating the problem of feature confusion between visually similar features and significantly improving device detection accuracy.

[0067] In addition, corresponding to Figure 1 In addition to the method shown in FIG. 1 , another embodiment of this embodiment further provides a data processing device for device detection. Figure 5 : is a schematic structural diagram of the data processing device 500, comprising: A model building module 510 is used to build an object detection model, wherein the object detection model includes an image encoder, a text encoder, a feature fuser, and a detector; a feature encoding module 5120 configured to encode a sample image of a target device using the image encoder to obtain image features of the target device; and to semantically encode a device description text of the target device using the text encoder to obtain semantic description features of the target device; A feature fusion module 530 is configured to perform bidirectional alignment encoding on the image features and the semantic description features through the feature fusion device to obtain enhanced image features and enhanced semantic description features; The model training module 540 is used to use the enhanced semantic description feature as the conditional input of the detector and train the detector based on the enhanced image feature and the label corresponding to the sample image; wherein the enhanced semantic description feature is used to generate the initial learnable parameters of the detector, and the detector generates a detection result by iterating the learnable parameters, and the detection result includes the detection box position and the device type; the label is marked with the device type of the target device and the detection box position corresponding to the target device in the sample image.

[0068] Optionally, the feature fusion module 530 performs bidirectional alignment encoding on the image features and the semantic description features through the feature fuser to obtain enhanced image features and enhanced semantic description features, including: executing through the feature fuser: encoding the image features based on a first cross-attention mechanism to obtain enhanced image features; encoding the semantic description features based on a second cross-attention mechanism to obtain enhanced semantic description features; wherein the key vector and value vector of the first cross-attention mechanism are determined based on the semantic description features, and the query vector is determined based on the image features; the key vector and value vector of the second cross-attention mechanism are determined based on the enhanced image features, and the query vector is determined based on the semantic description features.

[0069] Optionally, before respectively encoding the image features and semantic description features through the feature fuser, the feature fusion module 530 is also used to perform through the feature fuser: encoding the image features based on a deformable attention mechanism to obtain a first encoding result; encoding the semantic description features based on a self-attention mechanism to obtain a second encoding result; wherein the key vector and value vector of the first cross-attention mechanism are determined based on the second encoding result, and the query vector is determined based on the first encoding result; the key vector and value vector of the second cross-attention mechanism are determined based on the first encoding result, and the query vector is determined based on the second encoding result.

[0070] Optionally, the detector includes: multiple layers of decoder modules arranged in series, each decoder module performing: interacting the learnable parameters inputted by the current layer with the enhanced image features through a cross-attention mechanism to update the learnable parameters; performing a nonlinear transformation on the updated learnable parameters through a feedforward neural network to generate and output the iterated learnable parameters; wherein, in two adjacent decoder modules, the iterated learnable parameters outputted by the decoder module of the previous layer are directly used as the learnable parameters inputted by the decoder module of the next layer; a classification network predicting the device type based on the learnable parameters of the final iteration; and a regression network predicting the detection box position based on the learnable parameters of the final iteration.

[0071] Optionally, there are multiple sample images with corresponding sizes; the calculation formula of the deformable self-attention mechanism is:

[0072] in, is the size number corresponding to the sample image; is the number of attention heads; The number of neighboring features selected in each attention head; is the query vector; is the original position coordinate corresponding to the query vector; For the Image features of the sample image of a size; For the attention head, Size, query location, The offset corresponding to each adjacent feature; For the attention head, Size, query location, The attention weight of the neighboring features corresponding to the current position; No. The standardized atlas position after the normalization of the image features of each size; and All are The weight matrix corresponding to each attention point.

[0073] Optionally, the device description text includes device type information and brand information of the target device.

[0074] Optionally, the data processing device of this embodiment further includes: A model usage module is used to input the monitoring image and the device description text into the target detection model after training the detector to obtain a target detection result, wherein the target detection result includes the device type of the target device and the detection box position corresponding to the target device in the monitoring image.

[0075] It should be noted that the data processing device of this embodiment can be used as Figure 1 The execution subject of the method shown can thus realize Figure 1 The steps and functions in the method shown.

[0076] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 6 At the hardware level, the electronic device includes a processor and, optionally, an internal bus, a network interface, and memory. The memory may include internal memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for its services.

[0077] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 6 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0078] The memory is used to store the computer program. Specifically, the computer program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides the computer program to the processor.

[0079] Among them, the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming the above-mentioned Figure 5 The data processing device shown is applied to device detection. Correspondingly, the processor executes the program stored in the memory and is specifically used to perform the following operations: Building an object detection model, wherein the object detection model includes an image encoder, a text encoder, a feature fuser, and a detector; Encoding a sample image of a target device using the image encoder to obtain image features of the target device; and semantically encoding a device description text of the target device using the text encoder to obtain semantic description features of the target device; Performing bidirectional alignment encoding on the image features and the semantic description features by the feature fusion device to obtain enhanced image features and enhanced semantic description features; The enhanced semantic description feature is used as the conditional input of the detector, and the detector is trained based on the enhanced image feature and the label corresponding to the sample image; wherein the enhanced semantic description feature is used to generate the initial learnable parameters of the detector, and the detector generates a detection result by iterating the learnable parameters, and the detection result includes the detection box position and the device type; the label is marked with the device type of the target device and the detection box position corresponding to the target device in the sample image.

[0080] The above is as in this manual Figure 1The data processing methods disclosed in the illustrated embodiments can be applied to and implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be performed by hardware integrated logic circuits within the processor or by software instructions. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules within the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0081] Of course, in addition to software implementation, the electronic device in this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0082] In addition, an embodiment of the present application also proposes a computer-readable storage medium, which stores one or more computer programs, and the one or more computer programs include instructions.

[0083] When the above instructions are executed by a portable electronic device including multiple applications, the portable electronic device can execute Figure 1 The steps in the method shown include: An object detection model is constructed, wherein the object detection model includes an image encoder, a text encoder, a feature fuser, and a detector.

[0084] The sample image of the target device is encoded by the image encoder to obtain the image features of the target device; and the device description text of the target device is semantically encoded based on the text encoder to obtain the semantic description features of the target device.

[0085] The image features and the semantic description features are bidirectionally aligned and encoded by the feature fusion device to obtain enhanced image features and enhanced semantic description features.

[0086] The enhanced semantic description feature is used as the conditional input of the detector, and the detector is trained based on the enhanced image feature and the label corresponding to the sample image; wherein the enhanced semantic description feature is used to generate the initial learnable parameters of the detector, and the detector generates a detection result by iterating the learnable parameters, and the detection result includes the detection box position and the device type; the label is marked with the device type of the target device and the detection box position corresponding to the target device in the sample image.

[0087] Those skilled in the art will appreciate that the embodiments of this specification may provide methods, systems, or computer program products. Therefore, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0088] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0089] The above are merely examples of the present invention and are not intended to limit this specification. For those skilled in the art, various modifications and variations of this specification are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this specification shall be included within the scope of the claims of this specification. In addition, all other embodiments obtained by those of ordinary skill in the art without creative effort shall fall within the scope of protection of this document.

Claims

1. A data processing method for equipment detection, characterized in that: include: Building an object detection model, wherein the object detection model includes an image encoder, a text encoder, a feature fuser, and a detector; Encoding a sample image of a target device using the image encoder to obtain image features of the target device; and semantically encoding a device description text of the target device using the text encoder to obtain semantic description features of the target device; Performing bidirectional alignment encoding on the image features and the semantic description features by the feature fusion device to obtain enhanced image features and enhanced semantic description features; The enhanced semantic description feature is used as a conditional input of the detector, and the detector is trained based on the enhanced image feature and the label corresponding to the sample image; wherein the enhanced semantic description feature is used to generate initial learnable parameters of the detector, and the detector generates a detection result by iterating the learnable parameters, wherein the detection result includes the detection box position and the device type; The label is marked with the device type of the target device and the detection frame position corresponding to the target device in the sample image.

2. The method according to claim 1, characterized in that The image features and the semantic description features are bidirectionally aligned and encoded by the feature fusion device to obtain enhanced image features and enhanced semantic description features, including: The feature fuser performs: encoding the image features based on a first cross-attention mechanism to obtain enhanced image features; encoding the semantic description features based on a second cross-attention mechanism to obtain enhanced semantic description features; Among them, the key vector and value vector of the first cross-attention mechanism are determined based on the semantic description features, and the query vector is determined based on the image features; the key vector and value vector of the second cross-attention mechanism are determined based on the enhanced image features, and the query vector is determined based on the semantic description features.

3. The method according to claim 2, characterized in that Before encoding the image features and the semantic description features respectively by the feature fusion device, the method further includes: The feature fuser performs: encoding the image features based on a deformable attention mechanism to obtain a first encoding result; encoding the semantic description features based on a self-attention mechanism to obtain a second encoding result; Among them, the key vector and value vector of the first cross-attention mechanism are determined based on the second encoding result, and the query vector is determined based on the first encoding result; the key vector and value vector of the second cross-attention mechanism are determined based on the first encoding result, and the query vector is determined based on the second encoding result.

4. The method according to claim 1, wherein The detector comprises: Multiple layers of serially arranged decoder modules, each of which performs the following steps: interacting the learnable parameters input to the current layer with the enhanced image features through a cross-attention mechanism to update the learnable parameters; performing a nonlinear transformation on the updated learnable parameters through a feedforward neural network to generate and output the iterated learnable parameters; wherein, in two adjacent decoder modules, the iterated learnable parameters output by the decoder module of the previous layer are directly used as the learnable parameters input to the decoder module of the next layer; a classification network that predicts a device type based on the learnable parameters of a final iteration; The regression network predicts the detection box location based on the learnable parameters of the final iteration.

5. The method according to claim 1, wherein There are multiple sample images, and they have corresponding sizes. The calculation formula of the deformable self-attention mechanism is: in, is the size number corresponding to the sample image; is the number of attention heads; The number of neighboring features selected in each attention head; is the query vector; is the original position coordinate corresponding to the query vector; For the Image features of the sample image of a size; For the attention head, Size, query location, The offset corresponding to each adjacent feature; For the attention head, Size, query location, The attention weight of the neighboring features corresponding to the current position; No. The standardized atlas position after the normalization of the image features of each size; and All are The weight matrix corresponding to each attention point.

6. The method according to claim 1, characterized in that The device description text includes device type information and brand information of the target device.

7. The method according to claim 1, characterized in that After training the detector, the method further comprises: Obtain surveillance images of the target device; The monitoring image and the device description text are input into the target detection model to obtain a target detection result, wherein the target detection result includes a device type of the target device and a detection box position corresponding to the target device in the monitoring image.

8. A data processing device for equipment detection, characterized in that: include: A model building module is used to build an object detection model, wherein the object detection model includes an image encoder, a text encoder, a feature fuser and a detector; a feature encoding module, configured to encode a sample image of a target device using the image encoder to obtain image features of the target device; and to semantically encode a device description text of the target device based on the text encoder to obtain semantic description features of the target device; A feature fusion module, configured to perform bidirectional alignment encoding on the image features and the semantic description features through the feature fusion device to obtain enhanced image features and enhanced semantic description features; a model training module, configured to use the enhanced semantic description features as conditional inputs of the detector and train the detector based on the enhanced image features and labels corresponding to the sample images; wherein the enhanced semantic description features are used to generate initial learnable parameters of the detector, and the detector generates detection results by iterating the learnable parameters, wherein the detection results include detection box positions and device types; The label is marked with the device type of the target device and the detection frame position corresponding to the target device in the sample image.

9. An electronic device comprising: processor; and a memory arranged to store computer-executable instructions, wherein the executable instructions, when executed, cause the processor to perform the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer-readable storage medium storing a computer program, wherein: The computer program is operable to cause a computer to execute the method according to any one of claims 1 to 7.