Image recognition method and device, storage medium and electronic device

CN122799415APending Publication Date: 2026-09-22JIANGSU ZHONGTIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611307847.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-26
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0005]本申请实施例提供了一种图像识别方法和装置、存储介质及电子设备,以至少解决在少量标准样本条件下图像识别模型的识别性能较差的问题

Benefits of technology

[0011]采用本申请提供的上述实施例,首先根据包含多类别线缆字符的基础图像集合构建元学习任务,以在离线条件对初始特征提取网络进行训练,得到能够学习新类别线缆字符能力的元初始化模型参数;然后利用少量样本图像和合成样本图像构成的增强训练集,对通过加载元初始化模型参数所构建的初始识别模型进行迭代训练,不仅扩充了训练样本的数量且保证了样本的多样性,而且确保了训练后的图像识别模型能够准确识别新类别线缆字符。换言之,通过引入元学习算法,预先训练一个特征提取网络,使得利用训练得到的元初始化模型参数构建的初始识别模型能够具备新类别线缆字符的学习能力,同时也减少了利用增强训练集对初始识别模型的训练周期,解决了相关技术中少量标准样本条件下图像识别模型的识别性能较差的问题,提高了少量样本条件下的图像识别性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799415A_ABST
    Figure CN122799415A_ABST
Patent Text Reader

Abstract

The application discloses an image recognition method and device, a storage medium and an electronic device, comprising: based on a basic image set containing multiple categories of cable characters, constructing a meta-learning task, and based on the query set loss in the meta-learning task process, training an initial feature extraction network to obtain meta-initialization model parameters; acquiring a small number of sample images carrying category labels; constructing an initial recognition model by loading the meta-initialization model parameters, wherein the initial recognition model comprises a feature extraction network and a fully connected classifier; adjusting the initial recognition model using an enhanced training set composed of the small number of sample images and synthetic sample images to obtain a target recognition model used for recognizing the category of cable characters in real-time cable images, wherein the synthetic sample images are consistent with the image style of the small number of sample images. By adopting the technical scheme, the problem that the recognition performance of an image recognition model is poor under the condition of a small number of standard samples in the related art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to an image recognition method and apparatus, a storage medium, and an electronic device. Background Technology

[0002] During the manufacturing process of industrial cables or optical fibers, the cable surface is usually marked with character identifiers containing information such as model number, specifications, production date, and serial number. These identifiers are crucial for product traceability, inventory management, on-site construction, and operation and maintenance.

[0003] In related technologies, OCR or general deep learning models are mainly used to detect character markings on cable surfaces. However, these methods face many challenges in industrial scenarios. For example, in practical applications, due to frequent changes in product types on production lines, the font, size, arrangement, and printing process of character markings may differ for each new cable model. However, collecting and labeling hundreds or thousands of high-quality training images for each new cable specification is extremely costly and time-consuming, making it difficult to obtain a large number of labeled sample images. In this situation, the recognition performance of traditional methods drops significantly when there are insufficient samples, and overfitting is prone to occur, resulting in the technical problem of poor recognition performance of image recognition models under conditions of few standard samples.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This application provides an image recognition method and apparatus, storage medium and electronic device to at least solve the problem of poor recognition performance of image recognition models under conditions of a small number of standard samples.

[0006] According to one aspect of the embodiments of this application, an image recognition method is provided, comprising: constructing a meta-learning task based on a base image set containing multiple categories of cable characters, and training an initial feature extraction network based on the query set loss in the meta-learning task process to obtain meta-initialization model parameters capable of learning new categories of cable characters; acquiring a small number of sample images carrying category labels, wherein the small number of sample images includes character identifiers representing cable models and specifications; constructing an initial recognition model by loading the meta-initialization model parameters, wherein the initial recognition model includes a feature extraction network with the meta-initialization model parameters as initial weights and a fully connected classifier; adjusting the initial recognition model using an enhanced training set composed of the small number of sample images and synthetic sample images to obtain a target recognition model for recognizing cable character categories in real-time cable images, wherein the synthetic sample images have the same image style as the small number of sample images.

[0007] According to another aspect of the embodiments of this application, an image recognition device is provided, comprising: a first processing unit, configured to construct a meta-learning task based on a base image set containing multiple categories of cable characters, and to train an initial feature extraction network based on the query set loss during the meta-learning task process to obtain meta-initialization model parameters capable of learning new categories of cable characters; a first acquisition unit, configured to acquire a small number of sample images carrying category labels, wherein the small number of sample images includes character identifiers representing cable models and specifications; a second processing unit, configured to construct an initial recognition model by loading the meta-initialization model parameters, wherein the initial recognition model includes a feature extraction network with the meta-initialization model parameters as initial weights and a fully connected classifier; and an adjustment unit, configured to adjust the initial recognition model using an enhanced training set composed of the small number of sample images and synthetic sample images to obtain a target recognition model for recognizing cable character categories in real-time cable images, wherein the synthetic sample images have the same image style as the small number of sample images.

[0008] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the above-described image recognition method when running.

[0009] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0010] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to execute the image recognition method described above through the computer program.

[0011] Using the embodiments provided in this application, a meta-learning task is first constructed based on a base image set containing multiple categories of cable characters. This task is then used to train an initial feature extraction network offline, obtaining meta-initialized model parameters capable of learning new categories of cable characters. Next, an enhanced training set consisting of a small number of sample images and synthetic sample images is used to iteratively train the initial recognition model constructed by loading the meta-initialized model parameters. This not only expands the number of training samples and ensures sample diversity but also ensures that the trained image recognition model can accurately recognize new categories of cable characters. In other words, by introducing a meta-learning algorithm and pre-training a feature extraction network, the initial recognition model constructed using the trained meta-initialized model parameters can acquire the ability to learn new categories of cable characters. This also reduces the training cycle of the initial recognition model using the enhanced training set, solving the problem of poor recognition performance of image recognition models under conditions of few standard samples in related technologies and improving image recognition performance under conditions of few samples. Attached Figure Description

[0012] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and, together with the description thereof, serve to explain this application and do not constitute an undue limitation thereof. In the drawings:

[0013] Figure 1 This is a flowchart of an optional image recognition method according to an embodiment of this application.

[0014] Figure 2 This is an overall flowchart of an optional image recognition method according to an embodiment of this application.

[0015] Figure 3 This is a schematic diagram illustrating the task construction and learning of an optional meta-learning pre-training stage according to an embodiment of this application.

[0016] Figure 4 This is an optional online stage-controlled data generation and model adaptation flowchart according to an embodiment of this application.

[0017] Figure 5 This is an overall schematic diagram of the hardware architecture system used to execute the image recognition method in the embodiments of this application.

[0018] Figure 6 This is a schematic diagram of the structure of an optional image recognition device according to an embodiment of this application. Detailed Implementation

[0019] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0020] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0021] According to one aspect of the embodiments of this application, a method is provided as follows: Figure 1 The image recognition method shown includes:

[0022] Step S102: Based on the basic image set containing multiple categories of cable characters, a meta-learning task is constructed, and based on the query set loss in the meta-learning task process, the initial feature extraction network is trained to obtain meta-initialized model parameters capable of learning new categories of cable characters.

[0023] Step S104: Obtain a small number of sample images carrying category labels, wherein the small number of sample images include character identifiers representing cable models and specifications.

[0024] Step S106: By loading the meta-initialization model parameters, an initial recognition model is constructed, wherein the initial recognition model includes a feature extraction network with the meta-initialization model parameters as initial weights and a fully connected classifier.

[0025] Step S108: Using the enhanced training set composed of the small number of sample images and the synthetic sample images, the initial recognition model is adjusted to obtain a target recognition model for recognizing cable character categories in real-time cable images, wherein the synthetic sample images have the same image style as the small number of sample images.

[0026] To facilitate understanding of the above image recognition method, we will first combine... Figure 2The overall flowchart shown provides a brief overview.

[0027] like Figure 2 As shown, the technical solution proposed in this application is a few-sample cable character recognition method based on meta-learning and controllable data generation. Its core is to first use a large-scale base class dataset and combine it with a meta-learning algorithm to enable the initial feature extraction network to learn how to quickly learn the features of new categories of cable characters. Then, a small number or even a very small number of new samples are used to expand the sample images, thereby improving the recognition performance of the image recognition model for new categories of cable characters while significantly shortening the model training cycle.

[0028] Specifically, it includes two stages: the first stage is offline meta-learning training, and the second stage is online rapid adaptation and data generation with a small number of samples. The two stages are briefly described below.

[0029] The first stage aims to teach the model how to quickly learn new cable characters from a small number of samples. This is achieved by constructing a dataset containing a large number of known cable character categories (base classes) and training it using a model-independent meta-learning algorithm. Specifically, for each base class (i.e., each known cable character category), N-way and K-shot tasks are constructed. This involves randomly selecting N character categories, providing K support samples for each category (e.g., 3-way, 2-shot). This generates a large number of meta-learning tasks. The model is then processed through inner loops (task adaptation) and outer loops (meta-optimization) to rapidly improve the performance of the initial feature extraction network.

[0030] After training the initial feature extraction network through a large number of meta-learning tasks, a set of excellent initialization parameters (i.e., meta-initialization model parameters) is obtained. These parameters make the initial feature extraction network particularly sensitive to learning new tasks (i.e., new cable characters), and high performance can be achieved with only a small number of samples and a few gradient updates.

[0031] The second stage comprises two sub-stages. The first sub-stage, when dealing with a completely new type of cable character (a new category), involves acquiring only a very small number of labeled samples (e.g., fewer than 10 sample images), then loading a pre-trained feature extraction network and fine-tuning the last few layers of the network to adapt to the new category learning task. The second sub-stage involves globally fine-tuning the fine-tuned feature extraction network and the fully connected classifier. Specifically, in the second sub-stage, a large number of synthetic images are generated using a GAN (Generative Adversarial Network), and these synthetic images, along with a small number of real images, form an augmented training set. The fine-tuned feature extraction network and a fully connected classifier form an initial recognition model. This initial recognition model is then fine-tuned using the augmented training set to obtain a target recognition model for identifying the cable character category in real-time cable images. The specific implementation processes of the first and second sub-stages will be described below with reference to specific embodiments.

[0032] In this embodiment, a meta-learning task is first constructed based on a base image set containing multiple categories of cable characters. Then, based on the query set loss during the meta-learning task, an initial feature extraction network is trained to obtain meta-initialized model parameters capable of learning new categories of cable characters. This step corresponds to the offline meta-learning pre-training stage in the first stage described above. In this stage, the system constructs a base class dataset (e.g., 200 classes) containing hundreds of known cable character categories and employs a model-agnostic meta-learning algorithm, trained in an N-way K-shot task format. This allows the model to accurately identify new categories of cable characters after only a few updates. Essentially, this process trains the model to master the meta-learning ability to inductively classify rules from a very small number of examples.

[0033] The meta-learning task includes providing a preset number of support samples for each category of cable character images in a small number of sample images.

[0034] The aforementioned small set of sample images carrying category labels includes character identifiers representing cable models and specifications, i.e., new category sample images. When changing models on the production line, only 3-10 clear images of the new cables need to be manually collected and labeled with their character categories. This sample set serves as the support set for the new task and is the true basis for the model to adapt to the new categories.

[0035] The above-mentioned construction of the initial recognition model by loading the meta-initialization model parameters essentially includes two steps. The first step is to fine-tune the last few convolutional layers of the offline-trained feature extraction network by loading the meta-initialization model parameters. The second step is to connect a fully connected classifier after the fine-tuned feature extraction network to obtain the initial recognition model.

[0036] Then, using a generative adversarial network and a small number of sample images, a large number of synthetic images are generated. These synthetic and small sample images are then combined to form an enhanced training set. Using this enhanced training set, the initial recognition model is globally fine-tuned to obtain an adjusted target recognition model. Figure 2 As shown, the target recognition model is deployed to the detection terminal for online recognition of new types of cable characters. The detection terminal includes, but is not limited to, industrial production sites and edge detection or edge recognition devices that are directly used for cable character recognition and judgment.

[0037] The embodiments provided in this application first construct a meta-learning task based on a base image set containing multiple categories of cable characters. This task trains the initial feature extraction network offline, obtaining meta-initialized model parameters capable of learning new categories of cable characters. Then, using an enhanced training set consisting of a small number of sample images and synthetic sample images, the initial recognition model constructed by loading the meta-initialized model parameters is iteratively trained. This not only expands the number of training samples and ensures sample diversity but also ensures that the trained image recognition model can accurately recognize new categories of cable characters. In other words, by introducing a meta-learning algorithm and pre-training a feature extraction network, the initial recognition model constructed using the trained meta-initialized model parameters can acquire the ability to learn new categories of cable characters. It also reduces the training cycle of the initial recognition model using the enhanced training set, solving the problem of poor recognition performance of image recognition models under conditions of a small number of standard samples in related technologies, and improving image recognition performance under conditions of a small number of samples.

[0038] As an optional example, the above-described meta-learning task is constructed based on a base image set containing multiple categories of cable characters. The initial feature extraction network is trained based on the query set loss during the meta-learning task to obtain meta-initialization model parameters capable of learning new categories of cable characters. This includes: randomly sampling N categories of cable characters from the base image set and setting K support samples for each of the N categories to obtain a support set, where N and K are positive integers greater than or equal to 2; determining N prototype vectors corresponding to the N categories of cable characters based on the support set, where one of the N prototype vectors represents the standard position of a support sample of a category in the feature space; performing the meta-learning task based on real-time acquired cable character images and the N prototype vectors, and training the initial feature extraction network based on the query set loss.

[0039] In an optional embodiment, the offline meta-learning pre-training process in the first stage refers to the following steps S11-S13.

[0040] S11, Meta-learning task construction: In each iteration, an “N-way K-shot” task is randomly sampled from the base class dataset, that is, N character categories are randomly selected, and each category provides K support samples (e.g., 3-way 2-shot).

[0041] Where N and K are positive integers greater than or equal to 2. For example... Figure 3 As shown, suppose a meta-learning task is performed by randomly sampling from the base class dataset. This task involves three categories: new sample A, new sample B, and new sample C, with two support samples provided for each category. The base class dataset may contain a large number of sample images of category A or new sample A. The two support samples, A11 and A12, are obtained through filtering.

[0042] Using the same method, samples B11 and B21 are selected as support samples for category B, and samples C11 and C12 are selected as support samples for category C. This yields a support set.

[0043] S12 performs an inner loop (i.e., task adaptation) on the initial feature extraction network.

[0044] Specifically, the support set from the meta-learning task is input into a feature extraction network (such as a ResNet backbone network), and the prototype vector (the average value of the features of samples within the class) is calculated for each class, such as... Figure 3 The prototype A shown is the prototype vector of category A, prototype B is the prototype vector of category B, and prototype C is the prototype vector of category C. Then, a query set (new samples of N categories) is used to calculate the query set loss, and the model parameters are quickly updated on this task using gradient descent.

[0045] The prototype vector can be understood, but is not limited to, as the central feature or typical representative of each class of sample images, representing the standard position of this class of sample images in the feature space. Theoretically, image features of samples of the same class will cluster around this central feature or standard position, while prototypes of different classes will be far apart. Furthermore, the dimension of the prototype vector is the same as the feature dimension of the initial feature extraction network output.

[0046] It should be noted that the offline meta-learning pre-training stage ultimately yields a trained feature extraction network and a prototype vector-based classification mechanism. This feature extraction network maps images to a feature space where similar samples cluster together and dissimilar samples are spaced further apart.

[0047] S13, the performance of the feature extraction network after internal training and updating on the query set is used to calculate the meta-loss (i.e., query set loss).

[0048] The goal of this loss is to ensure that the feature extraction network achieves optimal performance on the new task after a few updates in the inner loop. Then, the initial parameters of the feature extraction network are updated based on the meta-loss to obtain the meta-initialized model parameters. The calculation process of the query set loss will be described in detail below with specific examples.

[0049] As an optional implementation, the above-mentioned meta-learning task is performed based on the real-time acquired cable character images and the N prototype vectors, and the initial feature extraction network is trained based on the query set loss. This includes: acquiring M batches of real-time acquired samples, and performing the following processing on the i-th batch of real-time acquired samples from the M batches, where M is a positive integer greater than or equal to 2, and i is a positive integer greater than or equal to 1 and less than or equal to M: If the category of the i-th batch of real-time acquired samples is the j-th category among the N categories, then sequentially determine the relationship between each acquired sample in the i-th batch of real-time acquired samples and the j-th prototype vector. The cosine distance between the samples is given, where j is a positive integer greater than or equal to 1 and less than or equal to N. Based on the cosine distance and a preset distance threshold, the classification result of the i-th batch of real-time collected samples is determined. Based on the classification result and the true category of the i-th batch of real-time collected samples, the i-th query set loss of the i-th meta-learning task is determined, where the i-th batch of real-time collected samples is a subset of the query set. If the value of the i-th query set loss does not meet the convergence condition, the structural parameters in the initial feature extraction network are adjusted until the convergence condition is met, and the structural parameters at the end of training are determined as the meta-initialization model parameters.

[0050] like Figure 3 As shown, after creating the meta-learning task, M batches of real-time sampled samples are acquired. The M batches of real-time sampled samples are of the same category, and these M batches of real-time sampled samples are sample images acquired online in real time during the industrial production process.

[0051] The following uses one batch of real-time samples from M batches of real-time samples as an example to describe the calculation method of the query set loss in each round of training. Figure 3 As shown, assuming the category of the i-th batch of real-time collected samples is category A, the prototype vector other than category A is determined as the j-th (here j=1) prototype vector, following the method in the above embodiment. Then, it is only necessary to calculate the cosine distance between the feature vector of each real-time collected sample in the i-th batch of real-time collected samples and the j-th prototype vector, and then compare the cosine distance with a preset distance threshold to determine the classification result of each collected sample in the i-th batch of real-time collected samples.

[0052] For example, if the cosine distance between the feature vector of a real-time sample and the j-th prototype vector is greater than a preset distance threshold, the determination result is 0, indicating that the predicted classification result is not belonging to category A; conversely, if the cosine distance is less than or equal to the preset distance threshold, the determination result is 1, indicating that the predicted classification result is that the real-time sample belongs to category A.

[0053] The predicted classification results of each sample in the i-th batch of real-time collected samples are obtained in the above manner. Then, based on the predicted classification results of the i-th batch and the true category of the i-th batch of real-time collected samples, the i-th query set loss of the i-th learning task is calculated.

[0054] If the loss value of the i-th query set satisfies the preset convergence condition, the structural parameters of the initial feature extraction network are adjusted until the convergence condition is met. Then, for the (i+1)-th batch of real-time acquired samples, the loss value of the (i+1)-th query set is calculated, and the structural parameters of the initial feature extraction network are adjusted based on the value of the loss value of the (i+1)-th query set. In other words, after the category prediction of each batch of real-time acquired samples is completed, the query set loss is calculated once based on the predicted classification results and the true category of that batch, and the structural parameters of the initial feature extraction network are adjusted based on the value of the loss value of the first batch. The structural parameters of the feature extraction network after adjustment based on the value of the last query set loss among the M query set losses are determined as the meta-initialized model parameters.

[0055] In another optional embodiment, the initial feature extraction network can be trained or adjusted in a different way: M query set losses are calculated using the same method as for the i-th query set loss described above. Then, the average of the M query set losses is calculated. If the average of the calculated query set losses does not meet the preset convergence condition, the structural parameters of the initial feature extraction network are adjusted until the convergence condition is met. The structural parameters at the end of training are then determined as the meta-initialized model parameters.

[0056] The purpose of introducing the meta-learning algorithm in this application embodiment is to learn different categories of cable characters in a large number of tasks during the offline stage, so that the initial feature extraction network can have the ability to generalize from one instance to another, thereby enabling the online stage to quickly and accurately identify new categories of characters that have never been seen before with only a small number of samples.

[0057] As another optional implementation, the above method further includes: in response to the switching of the inkjet content in the terminal device, detecting the acquired (i+1)th batch of real-time collected samples, wherein the inkjet content includes cable characters; when the detection result indicates that the character information of the (i+1)th batch of real-time collected samples is different from that of the (i)th batch of real-time collected samples, switching the j-th prototype vector to the (j+1)th prototype vector; sequentially determining the cosine distance between each collected sample in the (i+1)th batch of real-time collected samples and the (j+1)th prototype vector; determining the classification result of the (i+1)th batch of real-time collected samples based on the cosine distance and the preset distance threshold, and determining the (i+1)th query set loss of the (i+1)th meta-learning task based on the classification result and the true category of the (i+1)th batch of real-time collected samples; and adjusting the structural parameters in the initial feature extraction network when the value of the (i+1)th query set loss does not reach the convergence condition.

[0058] It's easy to understand that in real-world applications, the styles of characters on industrial cables are diverse, leading to frequent character category switching. After switching to a new category of cable characters, the corresponding coding content is edited and then sent to the coding machine (i.e., the terminal device).

[0059] In response to a received switching command for the inkjet printing content (such as cable characters or specifications), the inkjet printer's detection software detects a change in the real-time sample and will then switch accordingly. Figure 3 The calculation channel for the prototype vector is shown.

[0060] For example, suppose that after performing prediction classification, query set loss calculation, and model adjustment on the i-th (i=50) batch of M batches of real-time acquired samples in the above embodiment, the terminal device responds that the category of the (i+1)-th batch of real-time acquired samples has switched from category A to category B. Then, it will automatically switch to calculating the cosine distance between the (j+1)-th prototype vector (i.e., prototype vector B corresponding to category B) and the image feature vector of the (i+1)-th batch of real-time acquired samples, and obtain the prediction classification results of each sample in the (i+1)-th batch of real-time acquired samples in sequence. Then, the (i+1)-th query set loss is calculated, and if the value of the (i+1)-th query set loss does not reach the convergence condition, the structural parameters in the initial feature extraction network are adjusted.

[0061] In other words, the predicted classification result for each batch of real-time collected samples only needs to be determined by calculating the cosine distance with a prototype vector. Then, when the category of adjacent batches of real-time collected samples changes, it automatically switches to another prototype vector to perform distance calculation, classification result prediction, and query set loss calculation, etc.

[0062] This is because the real-time sample acquisition process is linked to the printer (i.e., the printing device). When the printed content changes, the detection template (i.e., the prototype vector) will automatically switch, such as from prototype vector A to prototype vector B.

[0063] Through the aforementioned automatic switching mechanism, seamless switching of detection templates (i.e., prototype vectors) can be completed within milliseconds. No downtime or manual intervention is required, allowing the production line to operate continuously, meeting the real-time and low-latency demands of industrial settings, and significantly improving production cycle time and overall equipment efficiency.

[0064] Meanwhile, through the semantically aware prototype (i.e. prototype vector) switching mechanism, it is ensured that each new character category is bound to its own exclusive, highly discriminative prototype vector, avoiding technical problems such as inaccurate prediction results and low recognition accuracy caused by template expiration.

[0065] As an optional example, the above method of constructing an initial recognition model by loading the meta-initialization model parameters includes: loading the feature extraction network with the meta-initialization model parameters as the initial weights; and connecting the fully connected classifier after the feature extraction network to obtain the initial recognition model.

[0066] The aforementioned feature extraction network, initialized with meta-initialized model parameters as initial weights, aims to reuse the rapid adaptive capabilities learned during the meta-learning phase. In the offline phase, through meta-training on numerous meta-learning tasks, it develops a high sensitivity to the visual patterns of cable characters. That is, after meta-training, the feature extraction network not only learns to distinguish the local structures of cable characters (such as stroke breaks and concave / convex edges), but also masters how to extract discriminative semantic features from preferred images under limited sample conditions.

[0067] If an untrained feature extraction network is used directly, even with generated data, it is difficult to achieve stable recognition with 5-10 samples. However, by directly loading the model parameters initialized with meta-initialization as initial weights, it is equivalent to giving the model a teacher with rich learning experience. This allows the model to be fine-tuned directly in a high-dimensional semantic space without having to learn the shape and texture patterns of cable characters from scratch, thus significantly reducing the risk of overfitting.

[0068] Connecting a fully connected classifier after the feature extraction network is crucial for achieving the transition from feature encoding to category decision-making. As a parameterized classification module, the meta-fully connected classifier's weight vector can be updated via gradient backpropagation, effectively capturing subtle differences in feature distribution between generated and real samples. This enables accurate modeling of synthetic images that are style-consistent but semantically invariant. By connecting the feature extraction network to the fully connected classifier, the system can maintain the strong generalization capability of feature representation provided by meta-learning while introducing a learnable decision boundary. This allows the model to dynamically adjust the classification threshold based on the distribution of hundreds of synthetic images and real samples in the enhanced training set, thus maintaining high-confidence recognition even in noisy, multi-interference industrial environments.

[0069] By loading a meta-initialized feature extraction network and connecting it to a fully connected classifier, flexible fine-tuning of parameterized classification can be achieved while maintaining the strong generalization ability of meta-learning. This improves the efficiency of the initial recognition model in adapting to small sample sizes and synthetic images. Simultaneously, the classification boundary can be dynamically optimized with the augmented training set, enabling high-precision and robust recognition with only a small number of real samples, while balancing deployment speed and recognition stability. This makes it suitable for the rapid production changeover requirements of cable models in industrial scenarios.

[0070] As an optional example, the above-described method of adjusting the initial recognition model using an enhanced training set composed of the small number of sample images and synthetic sample images to obtain a target recognition model for identifying cable character categories in real-time cable images includes: generating the synthetic sample images based on the small number of sample images and a generative adversarial network; obtaining a first loss value of a first loss function by inputting the enhanced training set into the initial recognition model, wherein the first loss function is used to minimize the error between the predicted classification result of the fully connected classifier and the true category label; adjusting the structural parameters of the high-level backbone network in the feature extraction network and the structural parameters in the fully connected classifier based on the first loss value, while freezing the structural parameters of the low-level backbone network in the feature extraction network to obtain an adjusted recognition model; and performing global fine-tuning on the adjusted recognition model using the meta-initialization model parameters as the regularization center to obtain the target recognition model.

[0071] like Figure 4As shown, to address the potential lack of diversity in a small number of samples for new categories, this embodiment integrates a conditional generative adversarial network (GAN) to expand the training samples. The GAN consists of two neural networks: a generator and a discriminator, which collaborate to optimize through an adversarial training mechanism. The generator receives random noise and attempts to generate realistic fake data (such as cable character images), while the discriminator learns to distinguish between real and generated data. They evolve continuously in this game, with the generator trying to deceive the discriminator, and the discriminator striving to identify fake samples. Ultimately, when the generator's output samples are indistinguishable from the real data distribution, the system reaches equilibrium. This achieves high-quality and realistic image data generation, particularly suitable for enhancing model generalization ability by synthesizing diverse and style-consistent images with very few labeled samples.

[0072] As an optional implementation, the above-mentioned method of generating the synthetic sample image based on the few sample images and the generative adversarial network includes: extracting image style feature vectors from the few sample images; obtaining a generated image by inputting a random noise vector, a character category label feature vector, and the image style feature vector into the generator of the generative adversarial network; and performing adversarial training by inputting the few sample images and the generated image into the discriminator of the generative adversarial network to obtain the synthetic sample image with the same image style as the few sample images.

[0073] In the process of generating synthetic sample images using generative adversarial networks, the generator generates cable region images corresponding to the character category based on random noise and character category labels. The generator G is composed of three parts to the input (i.e., generator conditional injection), which are combined through a conditional fusion layer, as shown in the following formula (1):

[0074]

[0075] (1)

[0076] in, It is a random noise vector. It is the embedding vector of the character category label. From a small sample support set The style features extracted from them It is a lightweight coding network.

[0077] In this embodiment, the generator's training not only relies on the offline base class dataset (i.e., the set of basic images), but also utilizes Z real samples of the new category to micro-guide the generator during the online adjustment phase. By calculating the similarity between the generated samples and the real samples in the feature space, it is ensured that the generated images closely resemble the real scene of the current new category of cables in terms of style, background, and lighting conditions, rather than being simple cable images.

[0078] After generating hundreds of high-quality synthetic images for each new category using a fine-tuned generator, the synthetic sample images are mixed with the original Z real images to form an enhanced training set or enhanced support set. The enhanced training set is then used to perform final fine-tuning (i.e., the second-stage online adjustment) on the model that has undergone meta-learning initialization. The fine-tuned model is then deployed to a recognition terminal (such as an industrial camera or handheld device) for online recognition of the new category of cables. For details, please refer to [reference needed]. Figure 4 As shown below, the overall implementation process of the second-stage online adjustment, the rapid adaptation and data generation with few samples in the first sub-stage, and the global fine-tuning in the second sub-stage are described in detail below with reference to specific embodiments.

[0079] As can be seen from the description of the above embodiments, the second stage includes two sub-stages. The first sub-stage is to obtain only a very small number of labeled samples and fine-tune the last few layers of the feature extraction network to adapt to the learning task of the new category when dealing with a completely new type of cable character (new category). The second sub-stage is to perform global fine-tuning on the fine-tuned feature extraction network and the fully connected classifier.

[0080] In a specific embodiment, assuming that when faced with a completely new type of cable character (new category), there are only a very small number (Z samples, Z<10) of labeled samples, the overall process of the second-stage online adjustment is as follows:

[0081] (1) Rapid adaptation based on prototype vectors.

[0082] For a small support set of the new category First, the prototype vector of each new category is calculated using the following formula (2):

[0083] (2)

[0084] Where Z is the number of samples for each new category, and C is the category index, where C is a positive integer greater than or equal to 1. The prototype vector is calculated by traversing from 1 to C.

[0085] (2) Load the pre-trained feature extraction network and fix most of its layers (i.e., fix the bottom backbone network), and only fine-tune the last few layers to adapt to the new task. Specifically, the feature adaptation of the first sub-stage is achieved through the following formula (3):

[0086] (3)

[0087] Where X is the input sample (feature vector / image tensor). For real labels, These are the parameters for the later layers of the feature extraction network.

[0088] Minimizing the cross-entropy loss function LCE (i.e., the first loss function) as shown in formula (3) above can, but is not limited to, reduce the error between the classifier's predicted classification result and the true label on the augmented dataset Senhanced. Essentially, this allows the later layers of the backbone network of the feature extraction network to adapt to new categories, while simultaneously freezing the backbone network to prevent overfitting.

[0089] It should be noted that in the first few stages, the frozen backbone network parameters do not participate in the update. The parameters that do participate in the update include the last few layers of the feature extraction network (i.e., the high-level backbone network or high-level feature extraction layers), as well as the structural parameters of the fully connected classifier. Freezing the structural parameters of the bottom-level backbone network aims to prevent the destruction of the general knowledge learned by the feature extraction network during the offline pre-training stage. Fine-tuning the high-level feature extraction layers and classifier aims to enable the initial recognition model to quickly adapt to the new task while reducing the risk of overfitting. After the adjustments in the first sub-stage, the adjusted initial model is obtained.

[0090] (3) Perform global fine-tuning on the feature extraction network and the fully connected classifier after the first sub-stage is completed to obtain the target recognition model.

[0091] By constructing an enhanced training set based on a small number of real samples and style-guided synthetic data, and fine-tuning the feature extraction network and classifier in stages, while introducing meta-learning parameter regularization, the technical solution of this application effectively suppresses the risk of overfitting under few sample conditions, improves the robustness of the model to changes in lighting, differences in surface texture and character deformation in real industrial scenarios, and achieves high model recognition accuracy with only a small number of labeled samples (such as 3 to 10 images), and efficiently completes the rapid adaptation and deployment of cable character models for new production lines.

[0092] As an optional implementation, the above-mentioned global fine-tuning of the adjusted recognition model using the meta-initialization model parameters as the regularization center to obtain the target recognition model includes: constructing a regularization term using the meta-initialization model parameters as the regularization center; constructing a second loss function based on the first loss function and the regularization term, wherein the second loss function is used to optimize the classification performance of the adjusted recognition model while ensuring that the structural parameters in the feature extraction network and the structural parameters in the fully connected classifier do not deviate excessively from the meta-initialization model parameters; and fine-tuning the structural parameters in the feature extraction network and the structural parameters in the fully connected classifier based on the second loss value of the second loss function to obtain the target recognition model.

[0093] The global fine-tuning process in the second sub-stage is as follows, specifically as shown in the following formula (4):

[0094] (4)

[0095] In formula (4), LCE is the cross-entropy loss function (i.e., the second loss function). The initial parameters obtained in the meta-learning stage (which can also be understood as the meta-initialized model parameters) are used to perform global fine-tuning of the adjusted recognition model in the second sub-stage using the above formula (4).

[0096] The formula for the second cross-entropy loss function is as follows: (5)

[0097] (5)

[0098] X is the input sample (feature vector / image tensor). For real labels, Here, C represents the predicted label / probability distribution of the model, and C is the class index. for Regularization coefficient (set to 0.01). for Regular terms, This is the enhanced training set.

[0099] During fine-tuning, the model parameters... Apply constraints to prevent them from deviating excessively from the initial parameters obtained during the meta-learning phase. This helps prevent overfitting and catastrophic forgetting.

[0100] Meanwhile, the objective minimized in the above formula (4) is the value of the second cross-entropy loss and The purpose of regularization is to improve classification performance by ensuring that the model's structural parameters do not deviate excessively from the initial parameters of meta-learning, thereby shortening the training cycle and reducing the number of iterations.

[0101] The aforementioned non-excessive deviation from the meta-initialized model parameters can also be understood as the structural parameters of the feature extraction network and the structural parameters of the fully connected classifier relative to the meta-initialized model parameters. If the norm difference does not exceed the preset regularization coefficient, the classification performance of the adjusted recognition model can be further improved.

[0102] After completing the global fine-tuning in the second sub-stage, the adjusted target recognition model can be applied to the recognition of cable character images acquired in real time during industrial production. Specifically, the prototype vector of the new category is calculated using Z support samples; for the image to be recognized (the cable character image acquired online in real time), its feature vector is extracted, and the cosine distance between it and the prototype vector of each new category is calculated. The classification result of each image to be recognized is obtained based on the cosine distance and a preset distance threshold.

[0103] As can be seen, the first sub-stage minimizes the first cross-entropy loss function; more precisely, it minimizes the error between the fully connected classifier's predicted classification result and the true class label on the augmented training set. Essentially, this allows the later layers of the feature extraction network to adapt to the new class while freezing the backbone network to prevent overfitting. The second sub-stage minimizes the second cross-entropy loss and... The purpose of regularization is to further improve classification performance while ensuring that the model's structural parameters do not deviate excessively from the initial parameters of meta-learning.

[0104] To gain a clearer understanding of the technical solution of this application, the following will be combined with... Figure 5 The overall architecture diagram shown further explains the above image recognition method.

[0105] like Figure 5 As shown, the complete system for performing the above image recognition method includes, but is not limited to, an image acquisition module, an edge computing unit, and a local model management platform.

[0106] The image acquisition module can, but is not limited to, acquiring real-time cable images generated in industrial processes, for example, through industrial cameras, encoders, and high-brightness linear array light sources.

[0107] The edge computing unit includes the deployed final recognition model (i.e., target recognition model), and uses this model to perform real-time inference on the real-time acquired cable images to obtain cable character content (i.e., classification results), coordinates, and other content.

[0108] It should be noted that the training and use of the recognition model are carried out simultaneously, such as... Figure 5 As shown, after the target recognition model deployed in the edge computing unit identifies the classification of the acquired images, the new sample labels of the real-time acquired cable images are sent to the local model management platform through the human-computer interaction interface. Then, the model update engine executes... Figure 4 The second-stage online adjustment process, as shown, is used to automatically execute the few-shot adaptation process and obtain an updated target recognition model for the new model (i.e., the new category) through global fine-tuning. The updated target recognition model is then distributed to the edge computing unit in real time to complete the deployment and use of the new model.

[0109] It should be noted that the local model management platform also includes a base class image set (i.e., base class data) stored in the database, a small number of new samples, and a meta-learning pre-trained model library. These three are used to perform the processing in the meta-learning pre-training stage, while the conditional GAN ​​model library is used to generate synthetic sample images based on a small number of new samples to construct an enhanced training set.

[0110] As can be seen from the above embodiments, in Figure 2 The document presents the two core phases of the technical solution in this application in a timeline format. The first phase is an offline, one-off model "general learning" phase; the second phase is an online, "rapid customization" phase for specific new cables, which incorporates generative data augmentation and ultimately outputs a deployable, specialized model.

[0111] exist Figure 3 The core mechanism of meta-learning is illustrated through a specific example (3 classes, 2 samples per class). Instead of learning to identify specific classes A, B, and C, the model learns a general ability through training on massive amounts of such random tasks: how to quickly build a classifier (prototype) using a small number of samples (support set) and accurately classify new samples (query set). The outer loop optimization objective is to optimize the initial parameters of the model. This makes it the best starting point for rapid learning. Mathematical modeling of the meta-learning stage (for few-shot classification): For each class in the meta-learning task, its prototype vector... The calculation is as follows (6):

[0112] (6)

[0113] in, This represents the set of samples belonging to class C in the support set (i.e., a small number of samples). The parameter is Feature extraction network, It is the input image. These are the corresponding category tags. The number of samples in category C (i.e., K in K-shot).

[0114] exist Figure 4 The core of this approach lies in the online phase adjustments. The key strategy is to use a very small number of real-world samples not only for fine-tuning the model but, more importantly, as "conditions" to guide the Generative Adversarial Network (GAN). By injecting style features extracted from real-world samples and using feature matching loss for constraints, it ensures that the hundreds or thousands of synthesized images generated are highly consistent with the current production environment in terms of non-semantic features such as texture, lighting, and background, thereby effectively improving the robustness of the model after fine-tuning.

[0115] exist Figure 5 The document describes the hardware and software system that supports the aforementioned image recognition method. This is a typical cloud-edge collaborative architecture. The edge side is responsible for high-performance, low-latency real-time recognition. The cloud (or local server) can be understood as the brain, responsible for managing all models and data, and automatically triggering a complex few-shot adaptation process when needed (such as when switching to a new model), generating a customized model and pushing it to the edge. The human-machine interface acts as a bridge, allowing operators to intervene (such as uploading new samples), forming a complete closed-loop process.

[0116] In a preferred embodiment, a few-sample cable character recognition method based on meta-learning and data generation is provided, such as... Figure 2 As shown, it includes the following steps S21 to S24.

[0117] S21, offline meta-learning pre-training.

[0118] S21-1 collects a base class dataset containing 200 different cable character categories, with approximately 50 images for each category.

[0119] S21-2, Construct a feature extraction network with ResNet-18 as the backbone.

[0120] S21-3 is trained using the Reptile meta-learning algorithm. Each training batch constructs 20 5-way, 5-shot tasks. After tens of thousands of iterations, a set of meta-learning initialization parameters is obtained. .

[0121] S22, adaptation to new online categories.

[0122] Suppose a new batch of ASS-GYTS-12B1.3 model cables arrives at the workshop, but only 5 clearly labeled sample images are available. (Loading...) As initial weights for the model.

[0123] S23, Generation of controllable data.

[0124] S23-1, Load a pre-trained conditional GAN ​​(trained on the base class dataset). Input 5 new samples and extract their style feature vectors.

[0125] S23-2 uses new class character labels (such as "Z", "T", "_", etc.) and noise vectors that incorporate the style of real samples as conditions to drive the generator to generate synthetic sample images. Through discriminator and feature matching loss constraints, it is ensured that the generated sample images have the texture, lighting and background characteristics of real samples.

[0126] S23-3, generate 200 composite images for each character category.

[0127] S24, rapid model fine-tuning and deployment.

[0128] S24-1, mix 5 real samples with 20 synthetic samples to form an enhanced training set.

[0129] S24-2 involves adding a fully connected classifier after the feature extraction network, fine-tuning the model end-to-end using an augmented training set, setting the learning rate to a low value (e.g., 1e-4), and training for 50 epochs.

[0130] S24-3 converts the fine-tuned model into an engine such as TensorRT and deploys it to an industrial camera system at the end of the production line to achieve 100% automatic detection and character recording of the new cable model.

[0131] In another preferred embodiment, this application also provides a system for implementing the above-described image recognition method, specifically including as follows: Figure 5 The following steps are shown.

[0132] (1) Image acquisition module: includes an industrial camera, a ring light source and a trigger sensor, used to acquire high-quality images when the cable is moving or stationary.

[0133] (2) Edge computing module: also known as edge computing unit, is embedded in GPU (Graphics Processing Unit) or AI acceleration chip and carries the recognition algorithm software in the embodiments of this application.

[0134] (3) Model Management Server: Stores meta-learning pre-trained models, base class datasets, and conditional GAN ​​models. It is responsible for receiving new samples, triggering the online adaptation process, generating and distributing new models to the edge computing module.

[0135] (4) Human-computer interaction interface: used to display recognition results and alarm information, and to provide a new sample upload interface.

[0136] Compared with traditional technical solutions, the technical solution of this application has at least the following beneficial effects:

[0137] (1) Extremely low sample requirements: Through meta-learning, the model has a strong ability to generalize from one instance to another, and only a few (3 to 10) labeled samples are needed to achieve practical accuracy, which greatly reduces the cost of data collection and labeling.

[0138] (2) Rapid deployment and adaptation: For new types of cables, the entire adaptation process (including data generation and model fine-tuning) can be completed in a short time (e.g., within 1 hour), meeting the actual needs of rapid changeover in production lines.

[0139] (3) Strong robustness: The controllable data generation mechanism can effectively simulate various changes in real scenes (such as uneven lighting, slight dirt, and angle tilt), enhancing the model's generalization ability in real complex environments.

[0140] (4) High recognition accuracy: Combining the excellent initialization and sample diversity brought by meta-learning, the recognition accuracy is greatly improved compared with the direct use of few-sample fine-tuning or traditional data augmentation methods.

[0141] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0142] According to another aspect of the embodiments of this application, an image recognition device is also provided, such as... Figure 6As shown, the device includes: a first processing unit 602, configured to construct a meta-learning task based on a base image set containing multiple categories of cable characters, and train an initial feature extraction network based on the query set loss during the meta-learning task process to obtain meta-initialization model parameters capable of learning new categories of cable characters; a first acquisition unit 604, configured to acquire a small number of sample images carrying category labels, wherein the small number of sample images includes character identifiers representing cable models and specifications; a second processing unit 606, configured to construct an initial recognition model by loading the meta-initialization model parameters, wherein the initial recognition model includes a feature extraction network with the meta-initialization model parameters as initial weights and a fully connected classifier; and an adjustment unit 608, configured to adjust the initial recognition model using an enhanced training set composed of the small number of sample images and synthetic sample images to obtain a target recognition model for recognizing cable character categories in real-time cable images, wherein the synthetic sample images have the same image style as the small number of sample images.

[0143] Optionally, the first processing unit 602 includes: a sampling module, configured to randomly sample N categories of cable characters from the base image set, and set K support samples for each of the N categories to obtain a support set, wherein N and K are positive integers greater than or equal to 2; a first processing module, configured to determine N prototype vectors corresponding to the N categories of cable characters based on the support set, wherein one of the N prototype vectors is used to characterize the standard position of a support sample of a category in the feature space; and a second processing module, configured to perform the meta-learning task based on the real-time acquired cable character images and the N prototype vectors, and train the initial feature extraction network based on the query set loss.

[0144] Optionally, the second processing module includes: a first processing submodule, configured to acquire M batches of real-time acquired samples, and perform the following processing on the i-th batch of real-time acquired samples in the M batches, where M is a positive integer greater than or equal to 2, and i is a positive integer greater than or equal to 1 and less than or equal to M: when the category of the i-th batch of real-time acquired samples is the j-th category among the N categories, sequentially determine the cosine distance between each acquired sample in the i-th batch of real-time acquired samples and the j-th prototype vector, where j is greater than or equal to 1 and less than or equal to N. The positive integer; based on the cosine distance and the preset distance threshold, the classification result of the i-th batch of real-time collected samples is determined, and based on the classification result and the true category of the i-th batch of real-time collected samples, the i-th query set loss of the i-th meta-learning task is determined, wherein the i-th batch of real-time collected samples is a subset of the query set; if the value of the i-th query set loss does not meet the convergence condition, the structural parameters in the initial feature extraction network are adjusted until the convergence condition is met, and the structural parameters at the end of training are determined as the meta-initialization model parameters.

[0145] Optionally, the above apparatus further includes: a detection unit, configured to detect the acquired (i+1)th batch of real-time collected samples in response to the switching of the inkjet content in the terminal device, wherein the inkjet content includes cable characters; a switching unit, configured to switch the j-th prototype vector to the (j+1)-th prototype vector when the detection result indicates that the character information of the (i+1)-th batch of real-time collected samples is different from that of the i-th batch of real-time collected samples; a third processing unit, configured to sequentially determine the cosine distance between each collected sample in the (i+1)-th batch of real-time collected samples and the (j+1)-th prototype vector; a fourth processing unit, configured to determine the classification result of the (i+1)-th batch of real-time collected samples based on the cosine distance and the preset distance threshold, and to determine the (i+1)-th query set loss of the (i+1)-th meta-learning task based on the classification result and the true category of the (i+1)-th batch of real-time collected samples; and a fifth processing unit, configured to adjust the structural parameters in the initial feature extraction network when the value of the (i+1)-th query set loss does not reach the convergence condition.

[0146] Optionally, the second processing unit 606 includes: a loading module for loading the feature extraction network with the meta-initialized model parameters as the initial weights; and a third processing module for connecting the fully connected classifier after the feature extraction network to obtain the initial recognition model.

[0147] Optionally, the adjustment unit 608 includes: a fourth processing module for generating the synthetic sample image based on the small number of sample images and the generative adversarial network; a fifth processing module for obtaining a first loss value of a first loss function by inputting the enhanced training set into the initial recognition model, wherein the first loss function is used to minimize the error between the predicted classification result of the fully connected classifier and the true category label; a first adjustment module for adjusting the structural parameters of the high-level backbone network in the feature extraction network and the structural parameters in the fully connected classifier based on the first loss value, while freezing the structural parameters of the low-level backbone network in the feature extraction network to obtain the adjusted recognition model; and a fine-tuning module for globally fine-tuning the adjusted recognition model with the meta-initialized model parameters as the regularization center to obtain the target recognition model.

[0148] Optionally, the fine-tuning module includes: a first construction submodule, used to construct a regularization term with the meta-initialized model parameters as the regularization center; a second construction submodule, used to construct a second loss function based on the first loss function and the regularization term, wherein the second loss function is used to optimize the classification performance of the adjusted recognition model while ensuring that the structural parameters in the feature extraction network and the structural parameters in the fully connected classifier do not deviate excessively from the meta-initialized model parameters; and a second processing submodule, used to fine-tune the structural parameters in the feature extraction network and the structural parameters in the fully connected classifier based on the second loss value of the second loss function to obtain the target recognition model.

[0149] Optionally, the fourth processing module includes: an extraction submodule for extracting image style feature vectors from the few sample images; a third processing submodule for obtaining a generated image by inputting a random noise vector, a character category label feature vector, and the image style feature vector into the generator of the generative adversarial network; and a fourth processing submodule for performing adversarial training by inputting the few sample images and the generated image into the discriminator of the generative adversarial network to obtain a synthetic sample image with an image style consistent with the few sample images.

[0150] By applying the aforementioned device to construct a meta-learning task based on a base image set containing multiple categories of cable characters, the initial feature extraction network is trained offline to obtain meta-initialized model parameters capable of learning new categories of cable characters. Then, using an enhanced training set consisting of a small number of sample images and synthetic sample images, the initial recognition model constructed by loading the meta-initialized model parameters is iteratively trained. This not only expands the number of training samples and ensures sample diversity but also ensures that the trained image recognition model can accurately recognize new categories of cable characters. In other words, by introducing a meta-learning algorithm and pre-training a feature extraction network, the initial recognition model constructed using the trained meta-initialized model parameters can possess the ability to learn new categories of cable characters. It also reduces the training cycle of the initial recognition model using the enhanced training set, solving the problem of poor recognition performance of image recognition models under conditions of few standard samples in related technologies and improving image recognition performance under conditions of few samples.

[0151] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0152] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. An image recognition method, characterized in that, include: Based on a base image set containing multiple categories of cable characters, a meta-learning task is constructed, and based on the query set loss in the meta-learning task process, the initial feature extraction network is trained to obtain meta-initialized model parameters capable of learning new categories of cable characters. Obtain a small number of sample images carrying category labels, wherein the small number of sample images include character identifiers representing cable type and specifications; An initial recognition model is constructed by loading the meta-initialization model parameters, wherein the initial recognition model includes a feature extraction network with the meta-initialization model parameters as initial weights and a fully connected classifier; The initial recognition model is adjusted using an enhanced training set consisting of the small number of sample images and synthetic sample images to obtain a target recognition model for recognizing cable character categories in real-time cable images, wherein the synthetic sample images have the same image style as the small number of sample images.

2. The method according to claim 1, characterized in that, The meta-learning task is constructed based on a base image set containing multiple categories of cable characters. Based on the query set loss during the meta-learning task, an initial feature extraction network is trained to obtain meta-initialized model parameters capable of learning new categories of cable characters, including: Randomly sample N categories of cable characters from the base image set, and set K support samples for each of the N categories to obtain a support set, where N and K are positive integers greater than or equal to 2; Based on the support set, N prototype vectors corresponding to the N categories of cable characters are determined, wherein one of the N prototype vectors is used to characterize the standard position of a support sample of a category in the feature space. Based on the real-time acquired cable character images and the N prototype vectors, the meta-learning task is performed, and the initial feature extraction network is trained based on the query set loss.

3. The method according to claim 2, characterized in that, The meta-learning task is performed based on the real-time acquired cable character images and the N prototype vectors, and the initial feature extraction network is trained based on the query set loss, including: Obtain M batches of real-time collected samples, and perform the following processing on the i-th batch of real-time collected samples in the M batches, where M is a positive integer greater than or equal to 2, and i is a positive integer greater than or equal to 1 and less than or equal to M: When the category of the i-th batch of real-time collected samples is the j-th category among the N categories, the cosine distance between each collected sample in the i-th batch of real-time collected samples and the j-th prototype vector is determined sequentially, where j is a positive integer greater than or equal to 1 and less than or equal to N; Based on the cosine distance and the preset distance threshold, the classification result of the i-th batch of real-time collected samples is determined, and based on the classification result and the true category of the i-th batch of real-time collected samples, the i-th query set loss of the i-th meta-learning task is determined, wherein the i-th batch of real-time collected samples is a subset of the query set; If the value of the loss of the i-th query set does not meet the convergence condition, the structural parameters in the initial feature extraction network are adjusted until the convergence condition is met, and the structural parameters at the end of training are determined as the meta-initialization model parameters.

4. The method according to claim 3, characterized in that, The method further includes: In response to the switching of the inkjet printing content in the terminal device, the acquired (i+1)th batch of real-time collected samples is detected, wherein the inkjet printing content includes cable characters; If the detection result indicates that the character information of the (i+1)th batch of real-time collected samples is different from that of the i-th batch of real-time collected samples, the j-th prototype vector is switched to the (j+1)-th prototype vector. The cosine distance between each sample in the (i+1)th batch of real-time collected samples and the (j+1)th prototype vector is determined sequentially. Based on the cosine distance and the preset distance threshold, the classification result of the (i+1)th batch of real-time collected samples is determined, and based on the classification result and the true category of the (i+1)th batch of real-time collected samples, the (i+1)th query set loss of the (i+1)th meta-learning task is determined. If the value of the loss of the (i+1)th query set does not meet the convergence condition, the structural parameters in the initial feature extraction network are adjusted.

5. The method according to claim 1, characterized in that, The step of constructing an initial recognition model by loading the meta-initialization model parameters includes: Load the feature extraction network with the meta-initialized model parameters as the initial weights; The fully connected classifier is connected after the feature extraction network to obtain the initial recognition model.

6. The method according to claim 1, characterized in that, The step of adjusting the initial recognition model using an enhanced training set composed of the small number of sample images and synthetic sample images to obtain a target recognition model for recognizing cable character categories in real-time cable images includes: The synthetic sample image is generated based on the small number of sample images and the generative adversarial network; By inputting the enhanced training set into the initial recognition model, a first loss value of the first loss function is obtained, wherein the first loss function is used to minimize the error between the predicted classification result of the fully connected classifier and the true class label; Based on the first loss value, the structural parameters of the high-level backbone network in the feature extraction network and the structural parameters in the fully connected classifier are adjusted, while the structural parameters of the low-level backbone network in the feature extraction network are frozen to obtain the adjusted recognition model. Using the meta-initialized model parameters as the regularization center, the adjusted recognition model is globally fine-tuned to obtain the target recognition model.

7. The method according to claim 6, characterized in that, The step of globally fine-tuning the adjusted recognition model using the meta-initialized model parameters as the regularization center to obtain the target recognition model includes: Using the aforementioned meta-initialized model parameters as the regularization center, a regularization term is constructed; Based on the first loss function and the regularization term, a second loss function is constructed, wherein the second loss function is used to optimize the classification performance of the adjusted recognition model while ensuring that the structural parameters in the feature extraction network and the structural parameters in the fully connected classifier do not deviate excessively from the meta-initialized model parameters; Based on the second loss value of the second loss function, the structural parameters in the feature extraction network and the structural parameters in the fully connected classifier are fine-tuned to obtain the target recognition model.

8. The method according to claim 6, characterized in that, The process of generating the synthetic sample image based on the small number of sample images and the generative adversarial network includes: Extract image style feature vectors from the small number of sample images; By inputting the random noise vector, the feature vector of the character category label, and the image style feature vector into the generator of the generative adversarial network, a generated image is obtained; By inputting the small number of sample images and the generated images into the discriminator of the generative adversarial network, adversarial training is performed to obtain the synthetic sample images that have the same image style as the small number of sample images.

9. An image recognition device, characterized in that, include: The first processing unit is used to construct a meta-learning task based on a base image set containing multiple categories of cable characters, and to train an initial feature extraction network based on the query set loss in the meta-learning task process to obtain meta-initialization model parameters that are capable of learning new categories of cable characters. The first acquisition unit is used to acquire a small number of sample images carrying category labels, wherein the small number of sample images include character identifiers representing cable models and specifications; The second processing unit is used to construct an initial recognition model by loading the meta-initialization model parameters, wherein the initial recognition model includes a feature extraction network with the meta-initialization model parameters as initial weights and a fully connected classifier; An adjustment unit is used to adjust the initial recognition model using an enhanced training set composed of the small number of sample images and synthetic sample images to obtain a target recognition model for recognizing cable character categories in real-time cable images, wherein the synthetic sample images have the same image style as the small number of sample images.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program can be executed by a terminal device or computer at runtime as described in any one of claims 1 to 8.

11. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 8.

12. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 8 through the computer program.