Equipment identification method and device based on spherical regularization and soft boundary ternary loss, electronic equipment and storage medium

By optimizing the device recognition method through spherical regularization and soft boundary ternary loss, the problems of diversity, category imbalance and training difficulty in device recognition technology are solved, and more efficient device recognition and robustness are achieved.

CN120635524APending Publication Date: 2025-09-12STATE GRID HUNAN ELECTRIC POWER COMPANY LIMITED +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510512343.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing device recognition technology faces problems such as diversity and complexity, category imbalance, and limitations in loss function design, which lead to difficulties in model training, insufficient recognition capabilities, and high resource consumption.

Method used

A method based on spherical regularization and soft boundary ternary loss is adopted to optimize the embedding space distribution by dynamically adjusting the boundary value and difficult sample mining strategy, and combined with model performance verification to improve training efficiency and recognition accuracy.

Benefits of technology

It enhances the adaptability and robustness of the model in diverse sample distributions, improves the ability to aggregate similar samples and separate heterogeneous samples, reduces the impact of class imbalance, and improves recognition accuracy and training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635524A_ABST
    Figure CN120635524A_ABST
Patent Text Reader

Abstract

The invention discloses an equipment identification method and device based on spherical regularization and soft boundary ternary loss, electronic equipment and a storage medium, and the method comprises the steps: S1, collecting an equipment image or video from an industrial scene, marking an equipment image sample, and generating a triple sample; s2, training a feature extraction network model based on spherical regularization and soft boundary ternary loss; s3, collecting an equipment image or video of an industrial scene in real time, and detecting an equipment position by using a pre-trained target detection model; and S4, extracting a device area, inputting a feature extraction network model, performing normalization operation to obtain embedded features, and matching the nearest category center according to the embedded features to complete classification. Through spherical regularization loss and soft boundary ternary loss optimization, distribution of equipment samples in an embedded space is improved, close aggregation of samples of the same kind and full separation of samples of different kinds are ensured, the influence of class imbalance on recognition is reduced, and recognition accuracy and robustness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of equipment monitoring technology, and in particular to an equipment identification method, apparatus, device and storage medium based on spherical regularization and soft boundary ternary loss. Background Art

[0002] In industrial scenarios, device recognition technology is of great significance for applications such as industrial automation, equipment monitoring, and industrial intrusion detection. However, the task of industrial device recognition faces the following challenges: Diversity and complexity: Industrial equipment is diverse, and the appearance characteristics of different devices vary significantly. Furthermore, due to environmental factors such as lighting, occlusion, shooting angle, and complex backgrounds, features in device images may be blurred or distorted, making the task of device recognition extremely complex. Class imbalance: Industrial equipment datasets often suffer from severe class imbalance. For example, some commonly used devices may have a large number of samples, while some rare devices may have fewer. This imbalance causes the model to favor classes with large sample sizes while neglecting classes with small sample sizes, affecting the comprehensiveness and robustness of recognition. Limitations of loss function design: Traditionally, triplet loss is a commonly used deep learning loss function that optimizes the embedding space by bringing similar samples closer together and dissimilar samples farther apart. However, setting a fixed margin can make training difficult: if the margin is too large, the model may overfit; if the margin is too small, the model may not fully learn the sample characteristics. The distance constraint imposed by fixed boundaries on samples cannot adapt to the diverse sample distributions in industrial scenarios. The embedding space design optimized by uniform distribution optimizes the separation between categories by constraining the uniformity of the distribution of embedded points: spherical regularization is used to constrain the distribution of embedded points on the unit sphere. Regularization is applied to the similarity between embedded points to enhance the uniformity of the distribution. Limitations: It is difficult to effectively deal with the problem of category imbalance. Common category features occupy the main area of ​​the spherical space, and the features of rare categories are marginalized, further exacerbating the category imbalance problem. The lack of focus on difficult samples leads to insufficient recognition of boundary samples. At the same time, the uniform distribution loss makes model training difficult, the training cycle is long, and it consumes huge computing resources.

[0003] It can be seen that the existing device identification technology has the following shortcomings:

[0004] Disadvantage 1: The existing ternary loss uses a fixed boundary value m, with max(0,m+d(z a ,z p )-d(z a ,z n)) is a constraint. Fixed boundary values ​​make it difficult to adapt to flexible, efficient, and high-performance model training. For some samples, excessively large boundary values ​​may cause training to struggle to converge; while too small a boundary value may lead to insufficient model discrimination. Fixed boundary values ​​limit the flexibility of the loss function, thereby reducing the model's adaptability and convergence.

[0005] Disadvantage 2: The difficult sample selection strategy and the loss function are not fully integrated, resulting in the mined samples failing to make the greatest contribution to the optimization of the model.

[0006] Disadvantage 3: The embedding features between categories are too concentrated, resulting in common categories occupying the main area of ​​the embedding feature space, and the embedding features of rare categories are marginalized, further exacerbating the category imbalance problem. Summary of the Invention

[0007] In response to the above technical problems, the present application provides a device identification method based on spherical regularization and soft boundary ternary loss.

[0008] This application is implemented through the following scheme:

[0009] A device identification method based on spherical regularization and soft boundary ternary loss comprises the steps of:

[0010] S1. Collect equipment images or videos from industrial scenes, annotate equipment image samples, and generate triplet samples;

[0011] S2. Train the feature extraction network model based on spherical regularization and soft boundary triple loss until the feature extraction network model achieves the expected indicators in the test set;

[0012] S3: Collect images or videos of equipment in industrial scenes in real time and use pre-trained object detection models to detect equipment locations.

[0013] S4. Extract the device area, input it into the feature extraction network model, perform normalization to obtain embedded features, match the nearest category center based on the embedded features, and complete the classification.

[0014] Furthermore, in step S1, the method of collecting device images or videos includes collecting by camera, drone inspection and robot inspection, and the collected data includes multiple states, viewing angles and lighting environments; the labeling of device image samples includes the category of the device; the triplet sample includes: anchor sample a, positive sample p, negative sample n; the anchor sample is any device image, the positive sample is a device image of the same type as the anchor sample, and the negative sample is a device image of a different type from the anchor sample.

[0015] Furthermore, the step S2 specifically includes the steps of:

[0016] S21. Sample input: Batch input of triplet samples (a, p, n) into the feature extraction network model, where the feature extraction network model can be any deep convolutional neural network.

[0017] S22. Calculate the embedding vector: The feature extraction network model extracts triplet image features (x a ,x p ,x n ), and then the image features are constrained to the unit sphere through normalization operation to obtain the embedded features (z a ,z p ,z n ), and at the same time, through the difficult sample mining strategy, we mine samples that are difficult to train, and use them for the later training to focus on the difficult samples;

[0018] S23. Calculate loss and optimize: Use the comprehensive loss L to update the feature extraction network model parameters;

[0019] S24. Performance testing: Verify whether the feature extraction network model achieves intra-class compactness and inter-class separation. If so, the training ends; otherwise, the training continues until intra-class compactness and inter-class separation are achieved.

[0020] Furthermore, the normalization operation is calculated as follows:

[0021]

[0022] Here, x refers to the device features obtained by the feature extraction network, and z refers to the embedded features obtained after the normalization operation of the device features.

[0023] Furthermore, the comprehensive loss L is composed of the soft boundary ternary loss L tri and spherical regularization loss L cos Composition, soft boundary ternary loss L tri is calculated as follows:

[0024]

[0025] Let the positive sample be close to the anchor point and the negative sample be far away from the anchor point, where d(,) represents the distance between the embedding vectors, L tri As d(z a ,z p )-d(z a ,z n ) Dynamic changes to accelerate the convergence of the model;

[0026] Spherical regularization loss L cos is calculated as follows:

[0027] L cos =cos(z a ,zn ) 2 +cos(z p ,z n ) 2

[0028] Intra-class samples are not subject to regularization constraints, and inter-class samples are distributed separately;

[0029] The comprehensive loss L is calculated as follows:

[0030] L=L tri +λ·L cos

[0031] Among them, λ is the balance factor, which balances the soft boundary ternary loss and the spherical regularization loss.

[0032] Furthermore, the difficult sample mining strategy is a negative sample that meets the following conditions:

[0033] d(z a ,z p )<d(z a ,z n )<d(z a ,z p )+η

[0034] Among them, η is the robust threshold of the distance between positive and negative samples and positive and positive samples.

[0035] Furthermore, the indicator for verifying whether the feature extraction network model achieves intra-class compactness and inter-class separation is the average intra-class distance d in and the average inter-class distance d out ; The criteria for intra-class compactness and inter-class separation are: the test set samples of each category meet, where β is the separation factor between the intra-class distance and the inter-class distance, and β=3.

[0036] Furthermore, in step S4, the category center is obtained by calculating the mean embedding feature of each category in the training data set, and is used as the representative vector of the category. The embedded feature of the device is matched with the category center, and the category label of the device is output.

[0037] On the other hand, the present application also provides a device identification apparatus based on spherical regularization and soft boundary ternary loss, comprising:

[0038] The triplet sample generation module is used to collect device images or videos from industrial scenes, annotate device image samples, and generate triplet samples;

[0039] The model training module is used to train the feature extraction network model based on spherical regularization and soft margin triplet loss until the feature extraction network model achieves the expected indicators in the test set;

[0040] The device detection module is used to collect images or videos of devices in industrial scenes in real time and detect the device location using a pre-trained object detection model;

[0041] The classification module is used to extract the device area, input the feature extraction network model, obtain the embedded features through normalization operation, match the nearest category center based on the embedded features, and complete the classification.

[0042] On the other hand, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the computer program, the steps of the device identification method based on spherical regularization and soft boundary ternary loss are implemented.

[0043] On the other hand, the present application also provides a storage medium, which includes a stored program, and when the program is run, controls the device where the storage medium is located to execute the steps of the device identification method based on spherical regularization and soft boundary ternary loss.

[0044] Compared with the existing technology, this application has the following beneficial effects:

[0045] 1. Loss function design with stronger dynamic adaptability

[0046] This application introduces a soft-margin ternary loss that dynamically adjusts the margin value based on a smooth logarithmic adaptive distance, avoiding the limitations of traditional fixed-margin ternary loss. This improves the model's adaptability to diverse sample distributions without requiring tedious hyperparameter tuning. It also enhances the model's convergence and stability, preventing overfitting and underfitting. This results in higher training efficiency and, in particular, superior robustness in complex industrial scenarios.

[0047] 2. The global optimization capability of the embedding space is stronger

[0048] This application designs a spherical distribution regularization loss to ensure the orthogonal distribution of embedding points on the unit sphere. At the same time, by optimizing the square of cosine similarity (excluding samples of the same type), it further enhances the discrimination of heterogeneous samples. This improves the distribution of rare categories and makes the features of small sample categories less likely to be marginalized. It also improves the global structural rationality of the embedding space and ensures that similar samples are clustered and heterogeneous samples are separated.

[0049] 3. Solve the optimization problem of difficult sample training

[0050] This application introduces a dynamic hard sample mining strategy, prioritizing hard samples with ambiguous features and boundaries for optimization. It then uses a combined loss consisting of a spherical regularization loss and a soft-margin ternary loss for training. This enhances the model's ability to learn complex samples and improves its ability to distinguish devices with similar features. This reduces interference from redundant samples in the training data, improving training efficiency and generalization.

[0051] 4. Model performance verification and efficient training design

[0052] This application uses a personalized design to verify the end of training through model performance, introduces intra-class distance and inter-class distance as indicators for model training termination, and dynamically detects the training effect (preferably the ratio of inter-class distance to intra-class distance) to ensure that the model achieves intra-class compactness and inter-class separation, avoids overtraining, saves resource overhead and improves efficiency.

[0053] In addition to the above-described purposes, features and advantages, the present application also has other purposes, features and advantages. The present application will be further described in detail below with reference to the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The drawings that constitute a part of this application are used to provide further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute improper limitations on this application.

[0055] Figure 1 It is a flow chart of a device identification method based on spherical regularization and soft boundary ternary loss in a preferred embodiment of the present application.

[0056] Figure 2 This is a schematic diagram of a device identification module based on spherical regularization and soft boundary ternary loss in a preferred embodiment of the present application.

[0057] Figure 3 This is a schematic block diagram of an electronic device entity according to a preferred embodiment of the present application.

[0058] Figure 4 It is a diagram of the internal structure of a computer device according to a preferred embodiment of the present application. DETAILED DESCRIPTION

[0059] The embodiments of the present application are described in detail below with reference to the accompanying drawings, but the present application can be implemented in a variety of different ways defined and covered below.

[0060] like Figure 1 As shown, a preferred embodiment of the present application provides a device identification method based on spherical regularization and soft boundary ternary loss, comprising the steps of:

[0061] S1. Collect equipment images or videos from industrial scenes, annotate equipment image samples, and generate triplet samples;

[0062] S2. Train the feature extraction network model based on spherical regularization and soft boundary triple loss until the feature extraction network model achieves the expected indicators in the test set;

[0063] S3: Collect images or videos of equipment in industrial scenes in real time and use pre-trained object detection models to detect equipment locations.

[0064] S4. Extract the device area, input it into the feature extraction network model, perform normalization to obtain embedded features, match the nearest category center based on the embedded features, and complete the classification.

[0065] To address the problems of the existing technology, this embodiment proposes a device recognition method based on spherical regularization and soft boundary ternary loss. This method improves the distribution of device samples in the embedding space through spherical regularization loss and soft boundary ternary loss optimization, ensures that similar samples are closely clustered and heterogeneous samples are fully separated, reduces the impact of class imbalance on recognition, and improves recognition accuracy and robustness. Among them, the soft boundary ternary loss uses a boundary with dynamic distance adaptive adjustment, avoids the problems caused by fixed margin through a smooth nonlinear function (log function), and improves the stability and convergence speed of model training. Combined with the loss function, it mines samples that are difficult to classify during training (samples with blurred boundaries between positive and negative samples or high feature similarity), focuses on learning difficult samples, improves the model's ability to distinguish complex samples, and significantly enhances the model's ability to recognize rare categories or boundary samples. The combination with the loss function can achieve accurate mining of difficult samples and rapid learning of the model for difficult samples. The pre-trained target detection model can be a YOLO series model. The device location is detected from the video screen or the image taken by the drone, and the screenshot of the device area corresponding to the location is input into the feature extraction network model in step S4. Because the pictures taken in reality do not only contain a single device, but may contain multiple devices and environmental backgrounds.

[0066] Specifically, in step S1, the method of collecting device images or videos includes collecting by camera, drone inspection and robot inspection, and the collected data includes multiple states, viewing angles and lighting environments; the labeling of device image samples includes the category of the device; the triplet sample includes: anchor sample a, positive sample p, negative sample n; the anchor sample is any device image, the positive sample is a device image of the same type as the anchor sample, and the negative sample is a device image of a different type from the anchor sample.

[0067] Specifically, the step S2 includes the following steps:

[0068] S21. Sample input: Batch input of triplet samples (a, p, n) into the feature extraction network model, where the feature extraction network model can be any deep convolutional neural network (such as ResNet, EfficientNet, etc.);

[0069] S22. Calculate the embedding vector: The feature extraction network model extracts triplet image features (x a ,x p ,x n ), and then the image features are constrained to the unit sphere through normalization operation to obtain the embedded features (z a ,z p ,z n ), and at the same time, through the difficult sample mining strategy, we mine samples that are difficult to train, and use them for the later training to focus on the difficult samples;

[0070] S23. Calculate loss and optimize: Use the comprehensive loss L to update the feature extraction network model parameters;

[0071] S24. Performance testing: Verify whether the feature extraction network model achieves intra-class compactness and inter-class separation. If so, the training ends; otherwise, the training continues until intra-class compactness and inter-class separation are achieved.

[0072] This embodiment adopts model performance verification and efficient training design when training the feature extraction network model based on spherical regularization and soft margin ternary loss: the personalized design of model performance verification ends the training, introduces intra-class distance and inter-class distance as indicators for terminating model training, and dynamically detects the training effect (preferably the ratio of inter-class distance to intra-class distance) to ensure that the model achieves intra-class compactness and inter-class separation, avoids overtraining, saves resource overhead and improves efficiency.

[0073] Specifically, the normalization operation is calculated as follows:

[0074]

[0075] Here, x refers to the device features obtained by the feature extraction network, and z refers to the embedded features obtained after the normalization operation of the device features.

[0076] Specifically, the comprehensive loss L is composed of the soft boundary ternary loss L tri and spherical regularization loss L cos Composition, soft boundary ternary loss L tri is calculated as follows:

[0077]

[0078] Let the positive sample be close to the anchor point and the negative sample be far away from the anchor point, where d(,) represents the distance between the embedding vectors, L tri As d(za ,z p )-d(z a ,z n ) Dynamic changes to accelerate the convergence of the model;

[0079] Spherical regularization loss L cos is calculated as follows:

[0080] L cos =cos(z a ,z n ) 2 +cos(z p ,z n ) 2

[0081] Intra-class samples are not subject to regularization constraints, and inter-class samples are distributed separately;

[0082] The comprehensive loss L is calculated as follows:

[0083] L=L tri +λ·L cos

[0084] Among them, λ is the balance factor, which balances the soft boundary ternary loss and the spherical regularization loss.

[0085] In this embodiment, L is composed of the soft boundary ternary loss L tri and spherical regularization loss L cos The comprehensive loss L has the following benefits: improving the distribution of device samples in the embedding space, ensuring that similar samples are closely clustered and heterogeneous samples are fully separated, reducing the impact of category imbalance on recognition, and improving recognition accuracy and robustness; at the same time, the soft boundary ternary loss can accelerate the convergence of the model.

[0086] Preferably, the difficult sample mining strategy is a negative sample that meets the following conditions:

[0087] d(z a ,z p )<d(z a ,z n )<d(z a ,z p )+η

[0088] Among them, η is the robust threshold of the distance between positive and negative samples and positive and positive samples.

[0089] In this embodiment, the difficult sample mining strategy is set to negative samples that meet the conditions of the above formula. Its benefits include: combining the loss function to mine samples that are difficult to classify during the training process (samples with blurred boundaries between positive and negative samples or high feature similarity), focusing on learning difficult samples, improving the model's ability to distinguish complex samples, and significantly enhancing the model's ability to recognize rare categories or boundary samples. The combination with the loss function can achieve accurate mining of difficult samples and rapid learning of difficult samples by the model.

[0090] Preferably, the indicator for verifying whether the feature extraction network model achieves intra-class compactness and inter-class separation is the average intra-class distance d in and the average inter-class distance d out ;

[0091] The criteria for intra-class compactness and inter-class separation are: the test set samples of each category meet d out >β·d in , where β is the separation factor between intra-class distance and inter-class distance, and β=3.

[0092] The benefits and purposes of setting corresponding indicators and standards in this embodiment include: ensuring that similar samples are closely clustered and heterogeneous samples are fully separated, reducing the impact of category imbalance on recognition, and improving recognition accuracy and robustness; accelerating model convergence and saving resource overhead for training the model.

[0093] Preferably, in step S4, the category center is obtained by calculating the mean embedding feature of each category in the training data set, and is used as the representative vector of the category. The embedded feature of the device is matched with the category center, and the category label of the device is output.

[0094] In this embodiment, the category center is obtained as the representative vector of the category by calculating the mean embedding feature for each category in the training data set. The embedded feature of the device is then matched with the category center, and the category label of the device is output. The benefits include: by matching the category center, the impact of abnormal samples and noise on the classification results is reduced, and the stability and robustness of the model are improved; at the same time, the category center has a clear physical meaning, which represents the central trend of the category, so the model output results are easier to understand and interpret.

[0095] like Figure 2 As shown, another preferred embodiment of the present application further provides a device identification apparatus based on spherical regularization and soft boundary ternary loss, comprising:

[0096] The triplet sample generation module is used to collect device images or videos from industrial scenes, annotate device image samples, and generate triplet samples;

[0097] The model training module is used to train the feature extraction network model based on spherical regularization and soft margin triplet loss until the feature extraction network model achieves the expected indicators in the test set;

[0098] The device detection module is used to collect images or videos of devices in industrial scenes in real time and detect the device location using a pre-trained object detection model;

[0099] The classification module is used to extract the device area, input the feature extraction network model, obtain the embedded features through normalization operation, match the nearest category center based on the embedded features, and complete the classification.

[0100] In summary, this application focuses on the combination of loss function design and model optimization strategy, as well as personalized design to improve training efficiency. The key technical points involved include:

[0101] 1. Innovative comprehensive loss design: a joint optimization mechanism of soft boundary ternary loss and spherical regularization loss.

[0102] A soft boundary ternary loss is designed, which uses a smooth nonlinear log function to achieve adaptive changes in the sample distance of the boundary, so that positive samples are infinitely close to the anchor point, and negative samples are infinitely far away from the anchor point. It overcomes the limitations of the fixed margin in the traditional ternary loss and is more distance-adaptive and easier to converge than the existing soft ternary loss. It significantly accelerates model convergence and improves stability and robustness.

[0103] A spherical regularization loss is designed, and the directions of heterogeneous samples are orthogonal (the angle approaches 90°) to improve the separation between classes globally; it ensures that the embedding points of similar samples are closely clustered, while heterogeneous samples are distributed more evenly, alleviating the problem of class imbalance.

[0104] 2. Propose a difficult sample mining strategy: Combined with the loss function, it mines samples that are difficult to classify during the training process (samples with blurred boundaries between positive and negative samples or high feature similarity), focuses on learning difficult samples, improves the model's ability to distinguish complex samples, and significantly enhances the model's ability to recognize rare categories or boundary samples. The combination with the loss function can achieve accurate mining of difficult samples and rapid learning of the model for difficult samples.

[0105] 3. Model performance verification and efficient training design: Model performance verification is personalized at the end of training. Intra-class distance and inter-class distance are introduced as indicators for model training termination. The training effect is dynamically detected (preferably the ratio of inter-class distance to intra-class distance) to ensure that the model achieves intra-class compactness and inter-class separation, avoid overtraining, save resources and improve efficiency.

[0106] 4. Adaptability and application innovation in industrial scenarios

[0107] It supports data collection in complex industrial scenarios, employing a variety of image acquisition methods (such as cameras, drones, and robotic inspections), and adapting to diverse lighting, occlusion, and background complexity. A comprehensive loss design combined with orthogonal distribution characteristics mitigates the negative impact of uneven class distribution on model training.

[0108] like Figure 3 As shown, a preferred embodiment of the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the device identification method based on spherical regularization and soft boundary ternary loss in the above embodiment are implemented.

[0109] like Figure 4 As shown, the preferred embodiment of the present application further provides a computer device, which can be a terminal or a liveness detection server, and its internal structure diagram can be as shown in FIG. Figure 4 As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with other external computer devices via a network connection. When the computer program is executed by the processor, the steps of the device identification method based on spherical regularization and soft boundary ternary loss are implemented.

[0110] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0111] A preferred embodiment of the present application also provides a storage medium, which includes a stored program, and when the program is run, controls the device where the storage medium is located to execute the steps of the device identification method based on spherical regularization and soft boundary ternary loss in the above embodiment.

[0112] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0113] If the functions described in the method of this embodiment are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a storage medium readable by one or more computing devices. Based on this understanding, the part of the embodiment of the present application that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computing device (which can be a personal computer, server, mobile computing device or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

[0114] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiment of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal translation scripting language JavaScript, etc.

[0115] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0116] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1The function specified in one or more boxes.

[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0118] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0119] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A device identification method based on spherical regularization and soft boundary ternary loss, characterized in that: Including steps: S1. Collect equipment images or videos from industrial scenes, annotate equipment image samples, and generate triplet samples; S2. Train the feature extraction network model based on spherical regularization and soft boundary triple loss until the feature extraction network model achieves the expected indicators in the test set; S3: Collect images or videos of equipment in industrial scenes in real time and use pre-trained object detection models to detect equipment locations. S4. Extract the device area, input it into the feature extraction network model, perform normalization to obtain embedded features, match the nearest category center based on the embedded features, and complete the classification.

2. The device identification method based on spherical regularization and soft boundary ternary loss according to claim 1, characterized in that: In step S1, the acquisition method of the image or video of the acquisition device includes camera, drone inspection and robot inspection acquisition, and the collected data includes multiple states, viewing angles and light environments; The device image samples are labeled including the device category; the triplet sample includes: anchor sample a, positive sample p, negative sample n; the anchor sample is any device image, the positive sample is a device image of the same type as the anchor sample, and the negative sample is a device image of a different type from the anchor sample.

3. The device identification method based on spherical regularization and soft boundary ternary loss according to claim 2, characterized in that: The step S2 specifically includes the following steps: S21. Sample input: Batch input of triplet samples (a, p, n) into the feature extraction network model, where the feature extraction network model can be any deep convolutional neural network. S22. Calculate the embedding vector: The feature extraction network model extracts triplet image features (x a ,x p ,x n ), and then the image features are constrained to the unit sphere through normalization operation to obtain the embedded features (z a ,z p ,z n ), and at the same time, through the difficult sample mining strategy, we mine samples that are difficult to train, and use them for the later training to focus on the difficult samples; S23. Calculate loss and optimize: Use the comprehensive loss L to update the feature extraction network model parameters; S24. Performance testing: Verify whether the feature extraction network model achieves intra-class compactness and inter-class separation. If so, the training ends; otherwise, the training continues until intra-class compactness and inter-class separation are achieved.

4. The device identification method based on spherical regularization and soft boundary ternary loss according to claim 3, characterized in that: The normalization operation is calculated as follows: Here, x refers to the device features obtained by the feature extraction network, and z refers to the embedded features obtained after the normalization operation of the device features.

5. The device identification method based on spherical regularization and soft boundary ternary loss according to claim 3, characterized in that: The comprehensive loss L is composed of the soft boundary ternary loss L tri and spherical regularization loss L cos Composition, soft boundary ternary loss L tri is calculated as follows: Let the positive sample be close to the anchor point and the negative sample be far away from the anchor point, where d(,) represents the distance between the embedding vectors, L tri As d(z a ,z p )-d(z a ,z n ) Dynamic changes to accelerate the convergence of the model; Spherical regularization loss L cos is calculated as follows: L cos =something a ,With n ) 2 +something(z p ,With n ) 2 Intra-class samples are not subject to regularization constraints, and inter-class samples are distributed separately; The comprehensive loss L is calculated as follows: L=L tri +λ·L cos Among them, λ is the balance factor, which balances the soft boundary ternary loss and the spherical regularization loss.

6. The device identification method based on spherical regularization and soft boundary ternary loss according to claim 3, characterized in that: The difficult sample mining strategy is a negative sample that meets the following conditions: d(z a ,z p )<d(z a ,z n )<d(z a ,z p )+η Among them, η is the robust threshold of the distance between positive and negative samples and positive and positive samples.

7. The device identification method based on spherical regularization and soft boundary ternary loss according to claim 3, characterized in that: The indicator for verifying whether the feature extraction network model achieves intra-class compactness and inter-class separation is the average intra-class distance d in and the average inter-class distance d out The criteria for intra-class compactness and inter-class separation are: the test set samples of each category meet d out >β·d in , where β is the separation factor between intra-class distance and inter-class distance, and β=3.

8. The device identification method based on spherical regularization and soft boundary ternary loss according to claim 1, characterized in that: In step S4, the category center is obtained by calculating the mean embedding feature of each category in the training data set, and is used as the representative vector of the category. The embedded feature of the device is matched with the category center, and the category label of the device is output.

9. A device identification apparatus based on spherical regularization and soft boundary ternary loss, characterized in that: include: The triplet sample generation module is used to collect device images or videos from industrial scenes, annotate device image samples, and generate triplet samples; The model training module is used to train the feature extraction network model based on spherical regularization and soft margin triplet loss until the feature extraction network model achieves the expected indicators in the test set; The device detection module is used to collect images or videos of devices in industrial scenes in real time and detect the device location using a pre-trained object detection model; The classification module is used to extract the device area, input the feature extraction network model, perform normalization operation to obtain embedded features, match the nearest category center based on the embedded features, and complete the classification.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the processor implements the steps of the device identification method based on spherical regularization and soft boundary ternary loss according to any one of claims 1 to 8.

11. A storage medium comprising a stored program, which controls the device where the storage medium is located to execute the steps of the device identification method based on spherical regularization and soft boundary ternary loss as described in any one of claims 1 to 8 when the program is run.

Citation Information

Cited By

  • Agricultural ecology intelligent monitoring regulation and control method and system based on machine learning

    CN121032144A

  • Model optimization method, electronic equipment and computer readable storage medium

    CN121561424A