Data processing method, data processing device, computer device, and computer program
The method improves image recognition accuracy and efficiency by fine-tuning models at field terminals and integrating server updates, addressing environmental adaptability and resource optimization challenges.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2024-06-03
- Publication Date
- 2026-05-19
AI Technical Summary
Existing image recognition models are not well adapted to various environments, leading to reduced learning performance and accuracy due to varying light intensity and background noise, and require significant computing power and time for training.
A data processing method involving fine-tuning of image recognition models at field terminals using local training sets and server updates, allowing models to adapt to specific environments and improve recognition accuracy while optimizing server resources.
Enhances model recognition accuracy and efficiency by enabling local fine-tuning and server optimization, making the model applicable to diverse environments and reducing computational load on servers.
Smart Images

Figure 2026515794000001_ABST
Abstract
Description
[Technical Field]
[0001] [Cross-reference of related applications] This application claims priority to Chinese Patent Application No. 202310894889.1, filed with the China National Intellectual Property Administration on July 20, 2023, with the title of the invention being "Data Processing Method, Related Apparatus, Device and Storage Medium," the entirety of which is incorporated herein by reference.
[0002] [Technical field] This application relates to the field of artificial intelligence technology, and more particularly to data processing methods, related apparatus, devices, and storage media. [Background technology]
[0003] In recent years, artificial intelligence (AI) technology has continued to develop, and its applications in the field of image recognition are expanding. By using complex algorithms and models, AI can recognize biometric information (faces, irises, palm prints, etc.), objects, and text within images, enabling intelligent image processing and analysis.
[0004] When collecting images in various environments, we often encounter complex environmental factors such as varying light intensity and background noise. These environmental factors can affect the accuracy of image recognition. Therefore, related technologies employ methods to train models by collecting large amounts of images in diverse environments in order to improve the model's recognition capabilities.
[0005] However, related technologies have a problem in that the images used to train models are not well adapted to various environments, limiting the types of samples the model can learn from, and consequently reducing the model's learning performance. On the other hand, training a model with a large number of images requires not only more computing power but also takes more time than expected. Effective solutions to these problems have not yet been found. [Overview of the project] [Means for solving the problem]
[0006] According to embodiments of this application, a data processing method, related apparatus, devices, and storage medium are provided that can make an image recognition model applicable to various specific field environments, improve the accuracy of model recognition, save server processing resources, and improve the learning efficiency of the model.
[0007] In view of this, according to one aspect of this application, a data processing method performed by a field terminal is provided. This method is The current field environment involves the steps of obtaining K (where K is an integer greater than or equal to 1) images by taking pictures with an image acquisition device, The steps include sending K images to a server, and the server obtaining K first prediction results using an image recognition model based on the K images, A step of constructing a fine-tuning training set according to K images and K first prediction results sent from a server, wherein the fine-tuning training set includes K groups of fine-tuning training data, and each group of fine-tuning training data includes an image and a first prediction result for that image. The process involves fine-tuning the model to be trained at the field terminal using a fine-tuning training set, obtaining a second prediction result corresponding to each image using the model to be trained based on the images included in the fine-tuning training data of each group within the fine-tuning training set, updating the model parameters of the model to be trained according to the second prediction result corresponding to each image and the first prediction result of the images in the fine-tuning training set, and obtaining a local recognition model and model adjustment parameters corresponding to the local recognition model. The process includes the steps of: if the local recognition model satisfies the model tuning conditions, sending model tuning parameters to a server, and the server updating the model parameters of the image recognition model according to a set of model tuning parameters from at least one terminal, wherein the set of model tuning parameters includes the model tuning parameters.
[0008] According to another aspect of this application, a data processing method performed by a server is provided. This method is The steps include receiving K (K is an integer greater than or equal to 1) images transmitted from a field terminal, wherein the K images were taken by the field terminal using a collection device under the current field environment, A step of obtaining K first prediction results using an image recognition model based on K images, The steps include: transmitting K first prediction results to a field terminal; constructing a fine-tuning training set according to K images and K first prediction results at the field terminal; fine-tuning the model to be trained at the field terminal using the fine-tuning training set; in the process of fine-tuning the model to be trained at the field terminal, obtaining a second prediction result corresponding to each image by the model to be trained based on the images included in the fine-tuning training data of each group within the fine-tuning training set; updating the model parameters of the model to be trained according to the second prediction result corresponding to each image and the first prediction results of the images in the fine-tuning training set; and obtaining a local recognition model and model tuning parameters corresponding to the local recognition model, wherein the fine-tuning training set includes K groups of fine-tuning training data, and each group of fine-tuning training data includes images and the first prediction results of the images. If the local recognition model satisfies the model fine-tuning conditions, the step is to receive model adjustment parameters sent from the field terminal, If a model tuning parameter set is obtained from at least one terminal, the step of updating the model parameters of the image recognition model, wherein the model tuning parameter set includes the model tuning parameters.
[0009] According to another aspect of this application, a data processing device is provided that is located at a field terminal. This device is A shooting module is configured to perform the step of obtaining K (where K is an integer greater than or equal to 1) images by taking pictures with an image acquisition device in the current field environment, A transmission module is configured to send the K images to a server, and for the server to perform the step of obtaining K first prediction results using an image recognition model based on the K images. An acquisition module is configured to perform the steps of constructing a fine-tuning training set according to the K images and the K first prediction results transmitted from the server, wherein the fine-tuning training set includes K groups of fine-tuning training data, and each group of fine-tuning training data includes an image and a first prediction result for the image. The update module is configured to perform the steps of: fine-tuning the model to be trained at the field terminal using the fine-tuning training set; in the process of fine-tuning the model to be trained at the field terminal, obtaining a second prediction result corresponding to each image using the model to be trained based on the images included in the fine-tuning training data of each group in the fine-tuning training set; updating the model parameters of the model to be trained according to the second prediction result corresponding to each image and the first prediction result of the images in the fine-tuning training set; and obtaining a local recognition model and model adjustment parameters corresponding to the local recognition model. The transmission module is further configured to transmit the model tuning parameters to the server if the local recognition model satisfies the model tuning conditions, and for the server to update the model parameters of the image recognition model according to a set of model tuning parameters from at least one terminal, wherein the set of model tuning parameters includes the model tuning parameters.
[0010] According to another aspect of this application, a data processing device located on a server is provided. This device is A receiving module is configured to perform the step of receiving K (K is an integer of 1 or more) images transmitted from a field terminal, wherein the K images are images taken by the field terminal with a collection device under the current field environment. An acquisition module configured to perform the step of obtaining K first prediction results using an image recognition model based on K images, A transmission module is configured to perform the following steps: transmit K first prediction results to a field terminal; the field terminal constructs a fine-tuning training set according to K images and K first prediction results; fine-tune the model to be trained at the field terminal using the fine-tuning training set; and in the process of fine-tuning the model to be trained at the field terminal, obtain a second prediction result corresponding to each image using the model to be trained based on the images included in the fine-tuning training data of each group within the fine-tuning training set; update the model parameters of the model to be trained according to the second prediction result corresponding to each image and the first prediction results of the images in the fine-tuning training set; and obtain a local recognition model and model tuning parameters corresponding to the local recognition model, wherein the fine-tuning training set includes K groups of fine-tuning training data, and each group of fine-tuning training data includes images and the first prediction results of the images. If a model tuning parameter set is obtained from at least one terminal, the update module is configured to perform the step of updating the model parameters of an image recognition model, wherein the model tuning parameter set includes model tuning parameters. The receiving module is further configured to perform the step of receiving model adjustment parameters transmitted from the field terminal if the local recognition model satisfies the model fine-tuning conditions.
[0011] According to another aspect of this application, a computer device is provided, comprising a memory in which a computer program is stored and a processor, wherein when the computer program is executed, the processor is made to implement the methods described in each aspect above.
[0012] According to another aspect of the present application, there is provided a computer-readable storage medium storing a computer program, which causes a processor to implement the method described in each of the above aspects when the computer program is executed.
[0013] According to another aspect of the present application, there is provided a computer program product including a computer program, which causes a processor to implement the method described in each of the above aspects when the computer program is executed.
Advantages of the Invention
[0014] As is clear from the above technical means, the embodiments of the present application have the following advantages.
[0015] In an embodiment of the present application, a data processing method is provided. First, in the current on-site environment, the on-site terminal captures K images using an image collection device. Since this on-site environment may affect the accuracy of image recognition, it is necessary to finely adjust the model. To achieve fine adjustment, the K images are sent to the server, and the server can obtain K first prediction results based on the K images using an image recognition model. Therefore, the first prediction results are used as labeling information for the corresponding images. Thus, the on-site terminal constructs a fine-tuning training set according to the K images and the K first prediction results sent from the server, and thereby can finely adjust the model to be trained at the on-site terminal using the fine-tuning training set. In the process of finely adjusting the model to be trained at the on-site terminal, the model to be trained can obtain a second prediction result corresponding to each image based on the images included in each group of fine-tuning training data in the fine-tuning training set. Therefore, the second prediction results are used as prediction information. Subsequently, the on-site terminal updates the model parameters of the model to be trained according to the second prediction results corresponding to each image and the first prediction results of the images in the fine-tuning training set, and obtains a local recognition model and model adjustment parameters corresponding to the local recognition model. When the local recognition model meets the model fine-tuning conditions, the on-site terminal sends the model adjustment parameters to the server, and the server updates the model parameters of the image recognition model according to the model adjustment parameter set from at least one terminal. By the above method, the terminal can finely adjust the local model based on the images collected in the on-site environment and report the finely adjusted target parameters to the server. The server updates the image recognition model by collectively using the parameter sets reported from each terminal. Thereby, the image recognition model can be applied to various specific on-site environments, not only improving the accuracy of model recognition, but also saving the processing resources of the server and improving the learning efficiency of the model.
Brief Description of the Drawings
[0016] [Figure 1] It is a schematic diagram of the implementation environment of the data processing method in an embodiment of the present application. [Figure 2] This is a schematic diagram of the implementation environment for an image recognition method according to one embodiment of this application. [Figure 3] This is a flowchart of a data processing method in one embodiment of the present application. [Figure 4] This is an interactive flowchart of voting based on a synchronous voting mechanism in one embodiment of the present application. [Figure 5] This is an interactive flowchart of voting based on a synchronous voting mechanism in another embodiment of the present application. [Figure 6] This is a schematic diagram for establishing the relationship between terminals in one embodiment of the present application. [Figure 7] This is a schematic diagram for establishing the relationship between terminals in another embodiment of the present application. [Figure 8] This is a schematic diagram for establishing the relationship between terminals in yet another embodiment of the present application. [Figure 9] This is an overall flowchart of a data processing method in one embodiment of this application. [Figure 10] This is a flowchart of a data processing method in one embodiment of the present application. [Figure 11] This is a framework diagram for performing data processing between a terminal and a server in one embodiment of the present application. [Figure 12] This is a schematic diagram of a data processing device according to one embodiment of the present application. [Figure 13] This is a schematic diagram of a data processing device in another embodiment of the present application. [Figure 14] This is a schematic diagram of a computer device in one embodiment of the present application. [Modes for carrying out the invention]
[0017] According to embodiments of this application, a data processing method, related apparatus, devices, and storage medium are provided that can make an image recognition model applicable to various specific field environments, improve the accuracy of model recognition, save server processing resources, and improve the learning efficiency of the model.
[0018] The terms “first,” “second,” “third,” “fourth,” etc. (if any) in the specification and claims of this application and in the drawings above are used to distinguish similar subjects and are not necessarily used to describe a particular order or sequence. The terms used herein are interchangeable under appropriate circumstances so that the embodiments of this application described herein may be carried out in an order other than that illustrated or described herein. Also, the terms “includes” and “corresponds” and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not necessarily limited to the steps or units explicitly listed, and may include other steps or units that are not explicitly listed or are specific to such process, method, product or device.
[0019] When collecting images, complex environmental factors are often encountered, and these factors can affect the accuracy of image recognition. Therefore, to improve the accuracy of image recognition, it is possible to train models with large amounts of training data and a large number of parameters. Large amounts of training data can provide sufficient model learning material for the model to learn, and a large number of parameters enhances the model's learning ability and makes it easier to learn the knowledge within the model learning material. However, training data is usually difficult to handle with all kinds of true environments, and the difficulty of training the model can increase with large amounts of training data and a large number of parameters.
[0020] Based on this, one embodiment of the present application provides a data processing method that can improve the effectiveness and stability of image recognition by fine-tuning the model, synchronizing the fine-tuning results, and optimizing the backend model according to images collected in real time at the site. The data processing method in the present application is applicable to at least one of the following scenarios.
[0021] 1. Biometric Authentication Scenario Biometric authentication technology refers to techniques that use computers in close conjunction with methods such as optics, sound, biosensors, and biostatistical principles to verify identity based on the unique biological characteristics of the human body (palms, face, iris, etc.). Because there are significant differences in image quality when images are collected in different environments, many challenges remain in improving the recognition accuracy of biometric authentication models. The following explanation will use palm recognition as an example.
[0022] Considering the complexity of field environments, including differences in light intensity and background noise—for example, dim lighting in a laboratory versus brighter lighting in an outdoor environment—recognition results may vary significantly across all palm scanning terminals if the same local recognition model with the same model parameters is used to recognize the collected images. Therefore, this application proposes that palm scanning terminals employ a local fine-tuning strategy to train the local recognition model, meaning that palm scanning terminals in different field environments can each employ a corresponding model optimization strategy. This improves the effectiveness and stability of palm authentication. Simultaneously, the palm scanning terminals also need to feed back the fine-tuned model tuning parameters to the server where the image recognition model is maintained. The server can then improve the recognition capabilities of the image recognition model by optimizing it based on the model tuning parameters reported from each palm scanning terminal.
[0023] Compared to the local recognition model of the palm scanning terminal, the server-side image recognition model has a larger number of model parameters and a relatively more complex model structure. Consequently, the image recognition model has stronger computational power and higher recognition accuracy. If a palm scanning terminal cannot recognize a palm image collected in the field using its local recognition model, the palm scanning terminal sends the palm image to the server. The server then calls the image recognition model to recognize the palm image and feeds the recognition result back to the palm scanning terminal to perform the corresponding task.
[0024] 2. Autonomous Driving Scenarios Image recognition is considered a crucial component of autonomous driving. Image recognition refers to the process of extracting features from images using computer technology for classification, recognition, and decision-making. In autonomous driving, image recognition primarily identifies various objects around the vehicle, such as pedestrians, road signs, and traffic lights, and plays a role in supporting autonomous driving, including decision-making.
[0025] Considering the complexity of the driving environment, including differences in weather, driving sections, and time of day, for example, light is weaker on cloudy or rainy days and stronger on sunny days; light is stronger when driving on elevated roads and weaker when driving through tunnels; and light is stronger during the day and weaker at night. If each in-vehicle terminal uses a local recognition model with the same model parameters to recognize the collected images, the recognition results may differ significantly. Therefore, in this application, the in-vehicle terminal employs a local fine-tuning strategy to train the local recognition model, meaning that in-vehicle terminals in different field environments can employ corresponding model tuning strategies. This improves the effectiveness and stability of object recognition. At the same time, the in-vehicle terminal also needs to feed back the fine-tuned model tuning parameters to the server side where a single image recognition model is maintained. The server can then improve the recognition capability of the image recognition model by optimizing it based on the model tuning parameters reported from each in-vehicle terminal.
[0026] If an in-vehicle terminal cannot recognize road images collected on-site using a local recognition model, the terminal sends these road images to a server. The server then calls an image recognition model to recognize the road images and feeds the recognition results back to the in-vehicle terminal, thereby enabling the vehicle to provide timely feedback.
[0027] 3. Security Measures Scenarios The security system is a complete system that transmits video signals in a closed loop using optical fiber, coaxial cable, or microwave, and can handle everything from capture to image display and recording. The security system not only significantly extends the observation distance of the human eye but also enhances the function of the human eye, allowing it to operate for extended periods in harsh environments, replacing manual labor.
[0028] Considering the complexity of real-world environments, including differences in weather, installation location, and time of day, if each security system uses a local recognition model with the same model parameters to recognize collected images, the recognition results may vary significantly. Therefore, in this application, the security system employs a local fine-tuning strategy to train the local recognition model, meaning that security systems in different field environments can employ corresponding model tuning strategies. This improves the effectiveness and stability of object recognition. At the same time, the security system also needs to feed back the fine-tuned model tuning parameters to the server side where a single image recognition model is maintained. The server can then improve the recognition capability of the image recognition model by optimizing it based on the model tuning parameters reported from each in-vehicle terminal.
[0029] If the security system cannot recognize images collected on-site using its local recognition model, it sends the images to a server. The server then invokes an image recognition model to recognize the images and feeds the recognition results back to the security system. If a security risk is detected, corresponding alarm information is triggered.
[0030] The above application scenarios are merely examples, and the data processing method according to this embodiment can be applied to other scenarios, but is not limited thereto.
[0031] It is understood that computer vision (CV) technology can be used to recognize images in this application. CV technology is the science of "making machines see," and more specifically, it refers to using machine vision—which uses cameras or computers to identify and detect objects in place of the human eye—and further processing these images with graphics to make them more suitable for human observation and instrument detection. As a science discipline, CV studies related theories and technologies with the aim of establishing artificial intelligence systems that can acquire information from images and multidimensional data. CV technology typically includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / motion recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning, and map construction, as well as common biometric authentication technologies such as facial recognition and fingerprint recognition.
[0032] The data processing method according to this application can be applied to the implementation environment shown in Figure 1. This implementation environment includes a field terminal 110 and a server 120, and the field terminal 110 and the server 120 are connected to each other via a communication network 130. Here, the communication network 130 uses standard communication technologies and / or protocols, and is usually the internet, but may be any network including, but not limited to, any combination of Bluetooth, local area network (LAN), metropolitan area network (MAN), wide area network (WAN), mobile, private network, or virtual private network. In some embodiments, customizable or dedicated data communication technologies may be used instead of, or in addition to, the data communication technologies described above.
[0033] Examples of field terminals 110 in this application include, but are not limited to, mobile phones, tablet computers, laptop computers, desktop computers, smart voice-interactive devices, smart home appliances, in-vehicle terminals, and aircraft. The client is implemented on the field terminal 110 and can run on the field terminal 110 in the form of a browser or as an independent application (app).
[0034] In this application, server 120 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), big data, and artificial intelligence (AI) platforms.
[0035] In combination with the above implementation environment, in step A1, the field terminal 110 transmits K images taken in the current field environment to the server 120 via the communication network 130. In step A2, the server 120 recognizes the K images and transmits a first prediction result for each image to the field terminal 110 via the communication network 130. In step A3, the field terminal 110 constructs a fine-tuning training set according to the K images and the K first prediction results transmitted from the server. In step A4, the field terminal 110 obtains K second prediction results using the model to be trained; that is, it obtains a second prediction result corresponding to each image using the model to be trained, based on each image included in the fine-tuning training data of each group in the fine-tuning training set. In step A5, the field terminal 110 trains the model to be trained according to the K second prediction results and the fine-tuning training set, and obtains a local recognition model and model tuning parameters. Specifically, the model to be trained is trained according to the second prediction result corresponding to each image and the first prediction result for that image in the fine-tuning training set, and the local recognition model and model tuning parameters are obtained. In step A6, the field terminal 110 transmits the model tuning parameters to the server 120 via the communication network 130. In step A7, the server 120 updates the model parameters of the image recognition model by combining the model tuning parameters reported from at least one terminal (i.e., the model tuning parameter set).
[0036] The implementation environment for the image recognition method will be described below, using the example that the field terminal 110 is a palm scanning terminal. Figure 2 is a schematic diagram of the implementation environment for the image recognition method in one embodiment of this application. Specifically, as shown in the figure, in step B1, the field terminal 110 recognizes the collected recognition image using a local recognition model and obtains a seventh prediction result. Here, the recognition image may be a palm image. In step B2, the seventh prediction result includes a category score. If the category score included in the seventh prediction result is equal to or greater than the category score threshold, the recognition image is determined to belong to the prediction category included in the seventh prediction result. In step B3, if the category score included in the seventh prediction result is less than the category score threshold, the field terminal 110 transmits the recognition image to the server 120 via the communication network 130. In step B4, the server 120 recognizes the recognition image using an image recognition model and obtains an image recognition result. In step B5, the server 120 transmits the image recognition results to the field terminal 110 via the communication network 130, which allows the field terminal 110 to perform the corresponding tasks according to the image recognition results.
[0037] In conjunction with the above explanation, the data processing method in this application will now be described from the perspective of the field terminal. Referring to Figure 3, the data processing method in the embodiment of this application may be completed independently at the field terminal, or it may be completed in cooperation with the field terminal and the server. The method of this application includes the following steps.
[0038] Step 210: Obtain K images by taking pictures with an image acquisition device under the current field conditions. Here, K is an integer greater than or equal to 1.
[0039] In one or more embodiments, the field terminal calls an image acquisition device (e.g., a camera, scanner, etc.) to take multiple images in the current environment, thereby obtaining K images.
[0040] Step 220: K images are sent to the server, and the server uses an image recognition model to obtain K first prediction results based on the K images.
[0041] In one or more embodiments, the field terminal can either send K images to the server sequentially, or directly package the K images and send them to the server in a batch. Based on this, the server inputs each image in the K images into an image recognition model, and the image recognition model outputs a first prediction result for each image, thereby obtaining K first prediction results. Here, each first prediction result includes the image's prediction category and category score.
[0042] The model described in this application is a deep learning model, and it is understandable that, for example, a convolutional neural network (CNN) is used. Deep learning is a machine learning technique that aims to simulate the workings of neurons in the human brain so that computers can learn autonomously and make decisions. Deep learning models are usually composed of multiple layers, and each layer can learn different levels of representation of the data.
[0043] In this application, the image recognition model deployed on the server side is a "large-scale model," meaning it has more powerful computing capabilities and higher recognition accuracy compared to the model deployed on the terminal side. It is trained on large amounts of data, learns a wider range of image features, and can accurately identify various objects. However, because of the enormous computational cost, it is usually deployed on the server side and is not suitable for execution on the terminal. In actual use, the local model on the terminal is often compared to and receives feedback from the "large-scale model" on the server in order to enable self-adjustment and optimization.
[0044] Step 230: Construct a fine-tuning training set according to K images and K first prediction results sent from the server. Here, the fine-tuning training set contains K groups of fine-tuning training data, and each group of fine-tuning training data contains an image and the first prediction result for that image.
[0045] In one or more embodiments, the field terminal can construct a fine-tuned training set according to the K images collected and the first prediction result for each image. The procedure for constructing a fine-tuned training set is described below using five images as an example.
[0046] Let's assume that K images sent from the field terminal to the server are Image 1, Image 2, Image 3, Image 4, and Image 5. When the server calls the image recognition model, it recognizes these images in order. Please refer to Table 1. Table 1 is a schematic diagram of the first prediction result for each recognized image. Here, we assume that the prediction category in the first prediction result is an object identifier, and that each object identifier uniquely represents one object (for example, User A).
[0047] [Table 1]
[0048] Based on this, we can construct fine-tuning training data for K groups. The fine-tuning training data for each group includes an image and the first prediction result for that image. For example, the fine-tuning training data for group 1 includes image 1, labeling category 10003, and labeling category score 0.95.
[0049] Step 240: The model to be trained on the field terminal is fine-tuned using the fine-tuning training set. During the process of fine-tuning the model to be trained on the field terminal, a second prediction result corresponding to each image is obtained by the model to be trained based on the images contained in the fine-tuning training data of each group within the fine-tuning training set.
[0050] In one or more embodiments, a fine-tuning training set is used to fine-tune the model to be trained at the field terminal. During the process of fine-tuning the model to be trained at the field terminal, the field terminal sequentially inputs the K collected images into the model to be trained, and the model outputs a second prediction result for each image. Here, each second prediction result includes the predicted category and category score of the image.
[0051] Step 250: Update the model parameters of the model to be trained according to the second prediction result corresponding to each image and the first prediction result of that image in the fine-tuning training set, and obtain the local recognition model and the model tuning parameters corresponding to the local recognition model.
[0052] In one or more embodiments, the fine-tuning training set includes K images and a first prediction result for each image. Here, the first prediction result is used as image labeling information, and a second prediction result for each image within the K images is used as image prediction information. Based on this, the model parameters of the model to be trained are updated using a corresponding loss function (e.g., a multiple classification loss function) based on the labeling and prediction information for each image, thereby obtaining a local recognition model and model tuning parameters corresponding to the local recognition model. Here, the model tuning parameters include, but are not limited to, model parameters, gradients, optimization algorithm parameters, and the fine-tuning training set.
[0053] In the embodiments of this application, updating the model parameters of the model to be trained can be understood as fine-tuning the model to be trained. In machine learning, fine-tuning is a transfer learning technique typically performed on a pre-trained model (e.g., a model trained on a large dataset). Typically, the parameters of the tuned model are fine-tuned based on a new, smaller dataset to optimize performance for a particular task.
[0054] Step 260: If the local recognition model satisfies the model tuning conditions, it sends the model tuning parameters to the server, and the server updates the model parameters of the image recognition model according to the model tuning parameter set from at least one terminal. Here, the model tuning parameter set includes the model tuning parameters.
[0055] In one or more embodiments, if the local recognition model satisfies the model fine-tuning conditions, the current fine-tuning method by the field terminal may be employed. Based on this, the field terminal can send model tuning parameters to the server. The server integrates the model tuning parameters reported from various terminals to obtain a model tuning parameter set. Based on this, the server uses the model tuning parameter set to update the model parameters of the image recognition model, i.e., fine-tune the image recognition model.
[0056] In embodiments of this application, a server can obtain an image recognition model with higher recognition efficiency based on distributed training. Here, distributed training refers to dividing the workload of the model to be trained and sharing it among multiple microprocessors (e.g., multiple terminals). Image recognition models have many parameters and a large amount of training data, exceeding the storage capacity of a single machine, so distributed parallel processing is necessary to speed up the model. Parallel mechanisms include data parallelism (DP), model parallelism (MP), pipeline parallelism (PP), and hybrid parallelism (HP). Structural designs include structures based on parameter servers, structures based on reduce, and structures based on message-passing interfaces (MPI).
[0057] Embodiments of this application provide a method for processing data. This method allows a terminal to fine-tune a local model based on images collected in the field environment and report the fine-tuned model tuning parameters to a server. The server integrates the model tuning parameter sets reported from each terminal to update the image recognition model. This makes the image recognition model applicable to various specific field environments, improving the accuracy of model recognition, saving server processing resources, and improving the efficiency of model learning.
[0058] In addition to one or more embodiments corresponding to Figure 3 above, another optional embodiment according to the embodiments of this application further includes: The steps include sending a model training request to a server, and the server determining the training dataset to apply to the field terminal according to the model training request. The initial training set is received from the server, where the initial training set includes initial training data for M groups, and the initial training data for each group includes images and the labeling results of those images. The initial recognition model obtains M initial prediction results based on M images included in the initial training set, where each initial prediction result includes the predicted category of the image and a category score. The step may include updating the model parameters of the initial recognition model according to the M labeled results and the M initial prediction results included in the initial training set, thereby obtaining a model to train.
[0059] In one or more embodiments, a method for obtaining the model to be trained is described. As is clear from the embodiments described above, after the field terminal is placed in the field, a model training request is sent from the field terminal to the server, and the server returns the training dataset to the field terminal. Note that other terminals can also generate the model to be trained in a similar manner, but this will not be explained in detail here.
[0060] Assuming the field terminal is a palm scanning terminal located within an amusement park, the server can obtain a training dataset from a large dataset based on the model training request sent from the field terminal. This training dataset includes palm images and image labeling results (such as user identifiers) of users registered for the amusement park's fast pass, and this training dataset is used as the initial training set. The field terminal receives the initial training set sent from the server. Here, the initial training set contains M groups of initial training data, and each group of initial training data contains an image and its labeling result. In other words, the initial training set contains M images.
[0061] Based on this, the field terminal sequentially inputs M images into the initial recognition model, and the initial recognition model outputs an initial prediction result corresponding to each image. Here, each initial prediction result includes the predicted category and category score of the image. Based on the labeling result and initial prediction result of each image, the model parameters of the initial recognition model can be updated using the corresponding loss function (e.g., a multiple classification loss function) to obtain a model to be trained.
[0062] The initial recognition model is a pre-training model (PTM), which is trained on a large amount of unlabeled data using a deep neural network (DNN) with a specific large number of parameters. The function approximation ability of the large-scale parameter DNN is used to enable the PTM to extract common features from the data, and techniques such as fine-tuning, parameter-efficient fine-tuning (PERT), and prompt-tuning are used to make it suitable for downstream tasks. Therefore, the PTM can achieve ideal results in few-shot or zero-shot scenarios.
[0063] Next, embodiments of this application provide a method for obtaining a model to be trained. By this method, the terminal obtains a locally usable model to be trained by training it using a training dataset delivered from a server. This enables the terminal to have image recognition capabilities, while also allowing the terminal to more effectively adapt to the local environment when fine-tuning the model.
[0064] In addition to one or more embodiments corresponding to Figure 3 above, another optional embodiment according to the embodiments of this application includes, before the step of obtaining K images by capturing them with an image acquisition device in the current field environment, The process involves obtaining on-site environmental information where the on-site terminal is located, where the on-site environmental information includes at least one of light intensity and background noise. If the light intensity included in the field environment information is not within the light intensity range, the first applied parameter of the image acquisition device is adjusted in response to a first adjustment operation on the image acquisition device, wherein the first applied parameter includes at least one of the shutter speed, sensitivity parameter and exposure compensation parameter. If the background noise included in the field environment information is greater than or equal to the background noise threshold, the second application parameter of the image acquisition device may be adjusted in response to a second adjustment operation on the image acquisition device, wherein the second application parameter includes at least one of a sharpness parameter, a sensitivity parameter, and a noise reduction parameter.
[0065] In one or more embodiments, a method for adjusting the image acquisition device is described. As is evident from the embodiments described above, it is also possible to adjust the image acquisition device based on field environment information in order to acquire higher quality images for use in training and inference of the model. The following describes this using a field terminal as an example. In actual use, other terminals can also be optimized in a similar manner, but this will not be explained in detail here.
[0066] In one feasible configuration, the field terminal's camera or light sensor can be used to acquire information about the field environment where the terminal is located. For example, the light intensity and background noise at the site can be acquired via the camera. Alternatively, the light intensity and color temperature at the site can be acquired via the light sensor.
[0067] For example, if the light intensity is not within the light intensity range (i.e., the light intensity is greater than or equal to the maximum light intensity, or less than or equal to the minimum light intensity), it is necessary to adjust the first applied parameters of the image acquisition device. Here, the first applied parameters include, but are not limited to, shutter speed, sensitivity parameter, and exposure compensation parameter. In one embodiment, if the light intensity is greater than or equal to the maximum light intensity, the shutter speed can be increased, the sensitivity (ISO) parameter can be decreased, and the exposure compensation (EV) parameter can be decreased. In another embodiment, if the light intensity is less than or equal to the minimum light intensity, the shutter speed can be decreased, the ISO parameter can be increased, and the EV parameter can be increased.
[0068] While it may be understandable to set the maximum light intensity to 1000 lux and the minimum light intensity to 10 lux, this is merely an example and should not be interpreted as limiting this application.
[0069] For example, if background noise exceeds the background noise threshold, it is necessary to adjust the second application parameter of the image acquisition device. Here, the second application parameter includes, but is not limited to, the sharpness parameter, ISO parameter, and noise reduction parameter. Based on this, if the background noise is high, the sharpness parameter can be lowered, the ISO parameter can be lowered, or the noise reduction parameter (e.g., spatial noise reduction parameter or temporal noise reduction parameter) can be increased.
[0070] While it may be understandable to set the background noise threshold to 50 decibels, this is merely an example and should not be interpreted as limiting this application.
[0071] In this application, the term "in response to" is used to indicate the conditions or states on which the execution of an operation depends, and one or more operations that can be performed when a particular condition or state is met. These operations may be performed in real time or with a certain delay.
[0072] Next, embodiments of this application provide a method for adjusting an image acquisition device. By adjusting the image acquisition device based on on-site environmental information as described above, the quality of the acquired images can be improved. Based on this, by automatically fine-tuning the local model, the local model on the terminal can be adapted to on-site illumination conditions, thereby improving the model's recognition capability.
[0073] In addition to the one or more embodiments corresponding to Figure 3 above, another optional embodiment according to the embodiments of this application updates the model parameters of the model to be trained according to the second prediction result corresponding to each image and the first prediction result of the images in the fine-tuning training set, and after the step of obtaining the local recognition model and the model tuning parameters corresponding to the local recognition model, further, The first step is to obtain the recognition accuracy of a local recognition model for N images, where N is an integer greater than or equal to 1, and the N images were taken by an image acquisition device. The step may include determining that the local recognition model satisfies the model fine-tuning condition if the recognition accuracy is above the accuracy threshold.
[0074] In one or more embodiments, a method for determining whether the model fine-tuning conditions are met is described. As is clear from the embodiments described above, after the model fine-tuning is completed at the field terminal, it is also necessary to evaluate the effect of the fine-tuning. That is, the recognition results of the local recognition model and the server-side image recognition model are compared. If the recognition result of the local recognition model matches or is close to the recognition result of the image recognition model, the fine-tuning is considered to have been successful. Conversely, it is necessary to re-execute the fine-tuning. The following explanation uses a field terminal as an example. In actual use, the quality of the fine-tuning results can be determined in a similar manner at other terminals, but this will not be explained in detail here.
[0075] The field terminal can use the collected field information to automatically fine-tune the local model to be trained, thereby obtaining a local recognition model. Here, fine-tuning refers to slightly adjusting the model parameters based on the model to be trained, according to new tasks or datasets. In this process, the fine-tuning operations are carried out according to the field environment information. For example, if the light intensity at the site changes, the model's sensitivity to changes in illuminance needs to be adjusted. Also, for example, if the background noise at the site increases, the model's noise resistance to interference may need to be enhanced. The fine-tuning process typically involves gradient descent and other optimization algorithms to minimize recognition errors under new environmental conditions.
[0076] One feasible form of this method for obtaining the recognition accuracy of a local recognition model for N images includes: transmitting N images captured by an image acquisition device to a server; the server obtaining N third prediction results using an image recognition model based on the N images; receiving the N third prediction results transmitted from the server; obtaining N fourth prediction results using the local recognition model based on the N images; and verifying the N fourth prediction results according to the N third prediction results to obtain the recognition accuracy for the N images.
[0077] In the actual execution process, the field terminal can send N images captured by the image acquisition device to the server. These N images may be those collected by the field terminal after training a local recognition model, or they may be a selection of images randomly chosen from K images. The N images satisfy the field illumination conditions of the field terminal. Based on this, the server uses each of the N images as input to the image recognition model, and the image recognition model obtains a third prediction result for each image. Here, the third prediction result includes the predicted category and category score for each image. Meanwhile, the field terminal uses each of the N images as input to the local recognition model, and the local recognition model obtains a fourth prediction result for each image. Here, the fourth prediction result includes the predicted category and category score for each image.
[0078] After receiving N third prediction results, the field terminal uses these N third prediction results as standard results to verify N fourth prediction results, thereby obtaining recognition accuracy for N images. For convenience of explanation, please refer to Table 2. Table 2 is a schematic diagram of the N third prediction results.
[0079] [Table 2]
[0080] Please refer to Table 3. Table 3 is a schematic diagram of the N fourth prediction results.
[0081] [Table 3]
[0082] Based on this, we compare N third prediction results with N fourth prediction results. If the third and fourth prediction results for an image match or are close to match, the recognition of that image is considered successful. We assume that the third and fourth prediction results being close means that "the prediction categories match and the absolute value of the difference between categories is 0.2 or less." Thus, the third and fourth prediction results for image 1 are close, the third and fourth prediction results for image 2 are different, the third prediction result for image 3 matches the fourth prediction result, the third and fourth prediction results for image 4 are different, and the third and fourth prediction results for image 5 are close. Therefore, 3 out of 5 images were recognized correctly, meaning the recognition accuracy is 0.6. If the recognition accuracy is above the accuracy threshold, the local recognition model of the field terminal satisfies the model fine-tuning conditions, meaning that the fine-tuning was successful.
[0083] In the method described above, N third prediction results output by the server via the image recognition model are used as criteria to verify N fourth prediction results. Since the image recognition model deployed on the server is relatively accurate, the third prediction results output from it are also relatively accurate and can be used as criteria, thus allowing for a more accurate verification of the recognition accuracy of the local recognition model.
[0084] Next, embodiments of this application provide a method for determining whether or not model fine-tuning conditions are met. This method enables the creation of a closed loop in which an on-site automatic fine-tuning module is used to acquire on-site environmental information, build a fine-tuning training set, fine-tune the model, and evaluate the fine-tuning results. As a result, the model can be automatically optimized to better adapt to on-site illumination conditions.
[0085] In addition to the one or more embodiments corresponding to Figure 3 above, another optional embodiment according to the embodiments of this application updates the model parameters of the model to be trained according to the second prediction result corresponding to each image and the first prediction result of the images in the fine-tuning training set, and after the step of obtaining the local recognition model and the model tuning parameters corresponding to the local recognition model, further, A step of obtaining the recognition accuracy of a local recognition model for N images, where N is an integer greater than or equal to 1, and the N images were taken by an image acquisition device. The process may include the following steps: if the recognition accuracy is equal to or greater than the accuracy threshold, send model tuning parameters to T terminals, each of the T terminals updates the model parameters of its corresponding model to be trained according to the model tuning parameters, and obtain T recognition models, where the T terminals are related to field terminals and T is an integer greater than or equal to 1; obtain a voting score corresponding to each terminal among the T terminals, where the voting score is determined according to the prediction results of the recognition model and the prediction results of the image recognition model; determine an overall recognition score according to the voting score corresponding to each terminal; and if the overall recognition score is equal to or greater than the recognition score threshold, determine that the local recognition model satisfies the model fine-tuning condition.
[0086] In one or more embodiments, an alternative method for determining whether the model fine-tuning conditions are met is introduced. As is evident from the embodiments described above, after the model fine-tuning is complete, the field terminal also needs to evaluate the effect of the fine-tuning using recognition accuracy. The method for calculating recognition accuracy can be found in the embodiments described above, so it will not be explained in detail here.
[0087] Specifically, if the recognition accuracy is above the accuracy threshold, each field terminal may send its model tuning parameters to T related terminals. This achieves the objective of synchronizing the model tuning parameters (i.e., the fine-tuning results), that is, synchronizing the model tuning parameters of a field terminal with other terminals in the same environment. Each terminal fine-tunes the model to be trained according to its own model tuning parameters and obtains its corresponding recognition model. Each terminal also needs to determine whether the model tuning parameters are appropriate or not based on a synchronized voting mechanism. If the model tuning parameters are appropriate, the terminal can cast a vote in favor; if the model tuning parameters are inappropriate, the terminal can cast a vote against. Exemplarily, the voting score corresponding to a vote in favor can be set to 1, and the voting score corresponding to a vote against can be set to 0. Based on this, an overall recognition score can be obtained using the voting scores of the T terminals. Whether the local recognition model satisfies the model fine-tuning conditions is determined using the overall recognition score.
[0088] In the embodiments of this application, the method for obtaining the recognition accuracy of the local recognition model for N images is the same as the method described above, and therefore will not be explained in detail here.
[0089] In this application, a synchronous voting mechanism is used to synchronize the results of model fine-tuning across multiple terminals. This mechanism allows terminals to vote based on the effectiveness of their model tuning parameters, deciding whether or not to accept the fine-tuning methods of other terminals. By aggregating the voting scores, a decision can be made as to whether or not to feed the fine-tuning methods back to the server for integration into the training of the image recognition model.
[0090] Next, embodiments of this application provide another method for determining whether the model fine-tuning conditions are met. By this method, the terminal optimizes the model and synchronizes the obtained model adjustment parameters with other terminals in the field, thereby further improving the accuracy of image recognition. Furthermore, by realizing a synchronized voting mechanism for fine-tuning results between terminals, the terminals can self-learn and optimize, thus improving the smartness of the system. Based on the synchronization and voting mechanism for fine-tuning results between terminals, optimization can be performed in real time according to field environment information, allowing the terminals to adapt more effectively to environmental changes and improving the image recognition effect.
[0091] In addition to one or more embodiments corresponding to Figure 3 above, in another optional embodiment according to the embodiments of this application, the step of determining an overall recognition score according to the voting score corresponding to each terminal is specifically: The steps include: aggregating the voting scores of T devices to obtain a total voting score; This could include a step to obtain an overall recognition score from the ratio of the total voting score to the T-value.
[0092] In one or more embodiments, a method for determining an overall recognition score based on voting scores is described. As is clear from the embodiments described above, the field terminal transmits model tuning parameters (i.e., fine-tuning results) to T terminals based on network communication technology. After each of the T terminals receives the model tuning parameters, it is necessary to fine-tune the local model to be trained according to the model tuning parameters and verify the effect of the fine-tuning. Verification methods include comparing the recognition accuracy of the model before and after the fine-tuning, or comparing the difference in recognition results between the fine-tuned recognition model and the server-side image recognition model.
[0093] Based on this, the terminals vote according to the fine-tuning effect, and majority rule is used in this process. This makes it possible to determine whether the current model tuning parameters can be adopted according to the voting results of T terminals. For ease of understanding, please refer to Figure 4. Figure 4 is an interactive flowchart of voting based on a synchronous voting mechanism in one embodiment of the present application. Specifically, as shown in the figure, it is as follows.
[0094] In step C1, the field terminal packages the model tuning parameters into a single information packet and sends it to terminal B. Here, the model tuning parameters include model parameters, gradients, and optimization algorithm parameters, where the optimization algorithm parameters include, but are not limited to, the optimization algorithm, learning rate, and number of iterations.
[0095] In step C2, the field terminal packages the model adjustment parameters into a single information packet and sends it to terminal C.
[0096] Note that the execution order of steps C1 and C2 is not restricted.
[0097] In step C3, terminal B fine-tunes the local model to be trained according to the received model tuning parameters to obtain a recognition model.
[0098] In step C4, terminal C fine-tunes the local model to be trained according to the received model tuning parameters to obtain a recognition model.
[0099] Note that the execution order of steps C3 and C4 is not restricted.
[0100] In step C5, terminal B performs a performance test based on the refined recognition model. If the model's performance improves after the refinement, a vote is cast in favor; conversely, if the model's performance deteriorates, a vote is cast against. For example, terminal B obtains recognition accuracy by comparing the prediction result output by the recognition model with the prediction result output by the image recognition model on the server. Next, terminal B votes according to the recognition accuracy.
[0101] In step C6, terminal C performs a performance test based on the refined recognition model. If the model's performance improves after the refinement, a vote is cast in favor; conversely, if the model's performance deteriorates, a vote is cast against. For example, terminal C obtains recognition accuracy by comparing the prediction result output by the recognition model with the prediction result output by the image recognition model on the server. Terminal C then votes according to the recognition accuracy.
[0102] Note that the execution order of steps C5 and C6 is not restricted.
[0103] In step C7, if the recognition accuracy obtained by terminal B is equal to or greater than the accuracy threshold, terminal B votes for the fine-tuning to be successful and receives a voting score of 1 point. Conversely, if the recognition accuracy obtained by terminal B is less than the accuracy threshold, terminal B votes for the fine-tuning to be unsuccessful and receives a voting score of 0 points.
[0104] In step C8, if the recognition accuracy obtained by terminal C is above the accuracy threshold, terminal C votes for the fine-tuning to be successful and receives a voting score of 1. Conversely, if the recognition accuracy obtained by terminal C is below the accuracy threshold, terminal C votes for the fine-tuning to be unsuccessful and receives a voting score of 0.
[0105] Note that the execution order of steps C7 and C8 is not restricted.
[0106] In step C9, as an example, one method is to have terminal B receive the voting score transmitted from terminal C and calculate an overall recognition score by combining it with its own voting score. Another example is for terminals B and C to each transmit their own voting scores to a field terminal, which then aggregates and calculates an overall recognition score. Next, the field terminal transmits this overall recognition score to terminal B.
[0107] In step C10, as an example, one method is to have terminal C receive the voting score sent from terminal B and calculate an overall recognition score by combining it with its own voting score. Another example is for terminals B and C to each send their own voting scores to a field terminal, which then aggregates and calculates an overall recognition score. Next, the field terminal sends this overall recognition score to terminal C.
[0108] Note that the execution order of steps C9 and C10 is not restricted.
[0109] In step C11, if the overall recognition score is equal to or greater than the score threshold (e.g., 0.5), it means that most terminals voted in favor. Therefore, terminal B accepts this fine-tuning method. If the overall recognition score is below the score threshold, it means that most terminals voted against it. Therefore, terminal B does not adopt this fine-tuning method.
[0110] In step C12, if the overall recognition score is equal to or greater than the score threshold (e.g., 0.5), it means that most terminals have voted in favor. Therefore, terminal C accepts this fine-tuning method. If the overall recognition score is below the score threshold, it means that most terminals have voted against it. Therefore, terminal C does not adopt this fine-tuning method.
[0111] The majority voting method will be explained below with reference to a specific example. Assuming that terminal B's voting score is 1 point and terminal C's voting score is 0 points, the total voting score will be 1 point (i.e., 1 + 0 = 1). The overall recognition score is calculated from the ratio of the total voting score to the T value. In this embodiment, T is 2. Based on this, the resulting overall recognition score is 0.5.
[0112] Furthermore, embodiments of this application provide a method for determining an overall recognition score based on voting scores. Using this method, each associated terminal votes on a fine-tuning result (i.e., a model tuning parameter), and the average voting score is calculated. This average is used as a criterion for deciding whether or not to accept the fine-tuning method. This provides a concrete approach for implementing the technical method, thereby improving its feasibility and operability.
[0113] In addition to one or more embodiments corresponding to Figure 3 above, in another optional embodiment according to the embodiments of this application, the step of determining an overall recognition score according to the voting score corresponding to each terminal is specifically: A step of obtaining a weighting parameter set corresponding to each terminal in T terminals, where the weighting parameter set includes at least one of device weighting, environment weighting, and priority weighting. For each of the T terminals, the voting score of the terminal is weighted using the terminal's weighting parameter set to obtain the terminal's weighted voting score. The procedure may include the step of determining an overall recognition score according to the voting weight score corresponding to each terminal among the T terminals.
[0114] In one or more embodiments, another method for determining an overall recognition score based on voting scores is introduced. As is evident from the embodiments described above, a field terminal transmits model tuning parameters (i.e., fine-tuning results) to T terminals based on network communication technology. After each of the T terminals receives the model tuning parameters, it is necessary to fine-tune the local model to be trained according to the model tuning parameters and verify the effect of the fine-tuning.
[0115] Based on this, terminals vote according to the fine-tuning effect, and a weighted voting system is employed in this process. Thus, it is possible to determine whether the current model tuning parameters can be adopted according to the voting results of T terminals. Here, in the weighted voting system, it is necessary to consider a set of weighted parameters corresponding to each terminal, and the set of weighted parameters includes at least one of device weighting, environment weighting, and priority weighting. For ease of understanding, please refer to Figure 5. Figure 5 is an interactive flowchart of voting based on a synchronous voting mechanism in another embodiment of the present application. Specifically, as shown in the figure, it is as follows:
[0116] Steps D1 to D8 in this embodiment are the same as steps C1 to C8 in the embodiment shown in Figure 4 above, so they will not be described in detail here.
[0117] In step D9, as an example, one method is given in which terminal B calculates a vote weighting score according to its own vote score and weighting parameter set. Another example is that terminal B sends its own vote score and weighting parameter set to a field terminal, and the field terminal calculates terminal B's vote weighting score according to the vote score and weighting parameter set sent from terminal B.
[0118] In step D10, as an example, one method is given in which terminal C calculates a vote weighting score according to its own voting score and weighting parameter set. Another example is that terminal C transmits its voting score and weighting parameter set to a field terminal, and the field terminal calculates terminal C's vote weighting score according to the voting score and weighting parameter set transmitted from terminal C.
[0119] Note that the execution order of steps D9 and D10 is not restricted.
[0120] In step D11, as an example, one method is to have terminal B receive the vote weighting score sent from terminal C and calculate an overall recognition score by combining its own weighting parameter set with that of terminal C. Another example is for terminals B and C to each send their own vote weighting scores to a field terminal, which then aggregates and calculates an overall recognition score. Next, the field terminal sends the overall recognition score to terminal B.
[0121] In step D12, as an example, one method is to have terminal C receive the vote weighting score sent from terminal B and calculate an overall recognition score by combining its own weighting parameter set with that of terminal B. Another example is for terminals B and C to each send their own vote weighting scores to a field terminal, which then aggregates and calculates an overall recognition score. Next, the field terminal sends the overall recognition score to terminal C.
[0122] Note that the execution order of steps D11 and D12 is not restricted.
[0123] Steps D13 to D14 in this embodiment are the same as steps C11 to C12 in the embodiment shown in Figure 4 above, and therefore will not be described in detail here.
[0124] The weighted voting method will be described below with reference to specific examples. Assume that the weighting parameter set includes device weighting, environment weighting, and priority weighting. Among these, the device weighting can be determined according to factors such as the performance of the terminal and its past prediction results. For example, the higher the performance of the terminal, the higher the device weighting. The environment weighting can be determined according to the on-site environment where the terminal is located. For example, when the light intensity in the on-site environment of the terminal is too strong or too weak, the environment weighting increases, while when the light intensity is within the light intensity range, the environment weighting decreases. The priority weighting can be determined according to the priority of the terminal for different optimization algorithm parameters. For example, when the optimization algorithm parameters used by the terminal match the optimization algorithm parameters of the on-site terminal, the priority weighting increases.
[0125] Here, it can be understood that the adjustment methods of device weighting, environment weighting, and priority weighting can be flexibly adjusted according to the actual situation. This is just an example and should not be construed as limiting the present application.
[0126] Based on this, the comprehensive recognition scores of terminal B and terminal C can be calculated according to the following formula.
[0127]
Equation
[0128] Furthermore, embodiments of this application provide another method for determining an overall recognition score based on voting scores. In this method, each associated terminal votes on a fine-tuning result (i.e., a model tuning parameter), and the voting scores are averaged by introducing weighting parameters for each terminal. This average is used as a criterion for deciding whether or not to accept the fine-tuning method. This not only provides a concrete approach for implementing the technical method, but also takes into account performance differences between different terminals, allowing for a more comprehensive vote on the fine-tuning results.
[0129] In addition to one or more embodiments corresponding to Figure 3 above, another implementable embodiment according to the embodiments of this application further includes the step of sending model adjustment parameters to T terminals before the step of sending the model adjustment parameters to T terminals. If a field terminal and at least one other terminal are located in the same area, the step is to determine that at least one other terminal is associated with the field terminal, and to determine that at least one other terminal is one of T terminals. or If the same binding object is set on the field terminal and at least one terminal, the step is to determine that at least one terminal is associated with the field terminal and to determine that at least one terminal is one of T terminals. or If a field terminal and at least one other terminal are connected to the same access point, the system may include the step of determining that at least one other terminal is associated with the field terminal and determining that at least one other terminal is one of T terminals.
[0130] In one or more embodiments, three methods for establishing association relationships between terminals are described. As is clear from the embodiments described above, a field terminal is associated with T terminals. Based on this, the field terminal can communicate with each of the T terminals. Three methods for establishing association relationships between terminals are described below with reference to specific examples.
[0131] Method 1: Based on geographical scope Specifically, devices belonging to the same geographical area are considered to be associated devices. This same geographical area can be an administrative district at the provincial, city, county, or township level, or it can be a custom area such as a residential area, school, or office building. Furthermore, devices have positioning capabilities, allowing users to establish relationships between devices located in the same geographical area.
[0132] For ease of understanding, please refer to Figure 6. Figure 6 is a schematic diagram for establishing the relationships between terminals in one embodiment of this application. As shown in the figure, if the residential area is considered as one location area, then terminals A, B, and C within the area of residential area A are related to each other. Terminals D, E, and F within the area of residential area B are also related to each other. However, each terminal within the area of residential area A is not related to each terminal within the area of residential area B. Therefore, if the field terminal is terminal A, then T terminals include terminals B and C.
[0133] Method 2: Based on binding relationships Specifically, users can customize the binding relationships between devices. For example, by assigning a merchant ID to each device, devices with the same merchant ID can be bound together.
[0134] For ease of understanding, please refer to Figure 7. Figure 7 is a schematic diagram for establishing the relationships between terminals in another embodiment of this application. As shown in the figure, we will use shopping mall A as an example. Assuming that shopping mall A has branches in cities A and B, terminals A, B, and C in shopping mall A in city A and terminals D, E, and F in shopping mall A in city B are interconnected. Therefore, if we consider terminal A as the on-site terminal, then T terminals include terminals B, C, D, E, and F.
[0135] Method 3: Based on network connectivity Specifically, an association relationship is established between devices connected to the same access point. Here, the access point may be a wireless hotspot [such as a mobile hotspot (wireless fidelity, WiFi)] or a wired access point.
[0136] For ease of understanding, please refer to Figure 8. Figure 8 is a schematic diagram for establishing the relationships between terminals in another embodiment of this application. As shown in the figure, terminals A, B, and C are connected to the same access point, and therefore the three terminals are interconnected. If terminal A is the field terminal, then the T terminals include terminals B and C.
[0137] Furthermore, embodiments of this application provide three methods for establishing correlation relationships between terminals. These methods establish correlation relationships between multiple terminals according to actual needs, allowing these terminals to perform corresponding processes (e.g., voting, status monitoring, etc.) as a single cluster, thereby improving the flexibility of implementing the technical method.
[0138] In addition to one or more embodiments corresponding to Figure 3 above, another optional embodiment according to the embodiments of this application further involves the step of determining whether the local recognition model satisfies the model fine-tuning conditions, Steps to obtain the image to be tested, The steps include obtaining a fifth prediction result using a local recognition model based on the image to be tested, The steps include obtaining T sixth prediction results from T terminals, where each sixth prediction result is obtained from the terminal by the recognition model based on the image to be tested, If, based on the fifth prediction result and the T sixth prediction results, it is determined that the local recognition model is in a model-stable state, the local recognition model may be used to perform the corresponding task.
[0139] In one or more embodiments, a method for monitoring the status of a model based on a plurality of associated terminals is described. As is evident from the embodiments described above, based on the field terminal and T terminals, the server can also monitor and evaluate the status of each terminal to determine whether these terminals have completed fine-tuning the model and whether the fine-tuned model is stable.
[0140] For example, monitoring and evaluating the status of a device primarily includes the fine-tuning status of the model and the operating status of the device. In other words, the device needs to be able to periodically send status reports to the server, and the report content will include the current fine-tuning status (e.g., fine-tuning in progress or completed) and the operating status (e.g., normal operation or malfunction). The server can understand the status of the device from these reports.
[0141] The stability of a model can usually be evaluated by comparing the model's recognition results for the same image (i.e., the image to be tested). For example, a field terminal uses the image to be tested as input to a local recognition model and obtains a fifth prediction result from the local recognition model. The other T terminals each use the image to be tested as input to a recognition model and obtain a sixth prediction result from the recognition model. If the fifth prediction result and the T sixth prediction results are the same or similar, the model is considered to be in a stable state. If the fifth prediction result and the T sixth prediction results differ significantly, further fine-tuning may be necessary.
[0142] In actual use, the server can compare the fifth prediction result with T sixth prediction results. Based on this, the server determines, depending on the terminal's status, whether the field terminal has completed fine-tuning and determines that the model is in a stable state according to the fifth prediction result and the T sixth prediction results. If both of these conditions are met, it is considered that the corresponding work has started at the field. Conversely, there may be cases where further fine-tuning of the model is necessary or the terminal needs maintenance.
[0143] Furthermore, if the prediction categories included in the fifth prediction result and the T sixth prediction results are the same, and the absolute value of the difference in category scores is less than or equal to a certain threshold (e.g., 0.2), then the prediction results are considered similar.
[0144] For ease of understanding, please refer to Figure 9. Figure 9 is an overall flowchart of a data processing method in one embodiment of this application. Specifically, as shown in the figure, it is as follows:
[0145] In step E1, the field terminal trains and deploys the model to be trained.
[0146] In step E2, the field terminal acquires field environment information using cameras, light sensors, etc.
[0147] In step E3, the field terminal uses an image acquisition device to capture K images, and uses a fine-tuning training set consisting of these K images to fine-tune the model to be trained on based on the current field environment information, thereby obtaining a local recognition model.
[0148] In step E4, the prediction results of the local recognition model are compared with the prediction results of the server-side image recognition model.
[0149] In step E5, if the prediction results of the local recognition model are the same as or similar to those of the image recognition model, the field terminal synchronizes the model tuning parameters to T terminals in the same environment.
[0150] In step E6, each of the T terminals in the same environment uses model tuning parameters to fine-tune the local model to be trained and obtain a recognition model.
[0151] In step E7, each of the T terminals votes in response to the performance changes of the fine-tuned model, and an overall recognition score is calculated based on the voting results. Depending on the magnitude of the overall recognition score, the T terminals decide whether to accept the fine-tuning. If so, step E8 is executed. Otherwise, step E9 is executed.
[0152] In step E8, if T terminals accept the adjustments, the field terminals report the model adjustment parameters to the server.
[0153] In step E9, if T devices do not accept this fine-tuning, the current model will remain unchanged.
[0154] In step E10, the server reruns the model training according to the model tuning parameters reported from the field terminal.
[0155] In step E11, the server monitors and evaluates the status of each terminal and the fine-tuned model performance.
[0156] In step E12, the server determines whether the model is currently in a stable state. If so, it executes step E13. Otherwise, it executes step E14.
[0157] In step E13, if the model is in a stable state, the terminal and server can perform their corresponding routine tasks.
[0158] In step E14, if the model is not in a stable state, the terminal and server will continue to fine-tune the model.
[0159] Furthermore, embodiments of this application provide a method for monitoring the status of a model based on a plurality of associated terminals. This method allows the plurality of associated terminals to monitor and evaluate the status of various devices in the field, and to determine whether these devices have completed fine-tuning the model and whether the fine-tuned model is stable. Based on this, the normal operation of the system can be maintained according to the determination. Additionally, because the status of the field terminals can be easily monitored in real time in the background, it contributes to the rapid detection and resolution of abnormal problems in the background.
[0160] In addition to one or more embodiments corresponding to Figure 3 above, another optional embodiment according to the embodiments of this application further involves the step of determining an overall recognition score according to the voting score corresponding to each terminal, If the local recognition model does not meet the model fine-tuning conditions, a model fine-tuning request is sent to T terminals within the terminal, and the terminals update the model parameters of the model to be trained according to the model fine-tuning request and obtain the recognition model, where T terminals are related to field terminals and T is an integer greater than or equal to 1. The steps include receiving model adjustment parameters sent from the terminal, This may include a step of updating the model parameters of the model to be trained on the field terminal using the model tuning parameters sent from the terminal.
[0161] In one or more embodiments, a training method is described for cases where the model fine-tuning conditions are not met. As is evident from the embodiments described above, if the local recognition model does not meet the model fine-tuning conditions, the field terminal can also randomly select one terminal from T terminals and send a model fine-tuning request to that terminal.
[0162] Assuming that "Terminal A" receives a model fine-tuning request, "Terminal A" can fine-tune the local model to be trained. Subsequently, "Terminal A" transmits the model tuning parameters obtained through the fine-tuning to other terminals within the same site, including the field terminal. Based on this, the field terminal fine-tunes the model to be trained using the model tuning parameters transmitted from "Terminal A". Alternatively, the field terminal continues to fine-tune the local recognition model using the model tuning parameters transmitted from "Terminal A".
[0163] Furthermore, embodiments of this application provide a training method for cases where the model fine-tuning conditions are not met. With the above method, if the local recognition model does not meet the model fine-tuning conditions, another terminal is selected from terminals within the same site to perform fine-tuning, and based on the model adjustment parameters obtained by fine-tuning this terminal, fine-tuning of other terminals within the same site can be continued, thereby enabling continuous model fine-tuning.
[0164] In addition to one or more embodiments corresponding to Figure 3 above, another optional embodiment according to the embodiments of this application updates the model parameters of the model to be trained according to K second prediction results and fine-tuning training sets, and after the step of obtaining a local recognition model and model tuning parameters corresponding to the local recognition model, further, If the local recognition model satisfies the model fine-tuning conditions, the step is to capture an image for recognition using an image acquisition device, Based on the recognition image, a local recognition model obtains a seventh prediction result, where the seventh prediction result includes a prediction category and a category score. The process may include a step of determining that the recognition image belongs to a predicted category in the seventh prediction result if the category score in the seventh prediction result is equal to or greater than the category score threshold.
[0165] In one or more embodiments, a method for performing image recognition locally on a terminal is described. As is clear from the embodiments described above, if the local recognition model satisfies the model fine-tuning conditions, the field terminal can use this local recognition model to perform image recognition. The following explanation uses a field terminal as an example. In actual use, other terminals can also perform image recognition in a similar manner, but this will not be explained in detail here.
[0166] Specifically, the field terminal captures multiple images using an image acquisition device and selects one high-resolution image from among them as the recognition image. Taking palm image recognition as an example, the collected multiple images undergo image quality evaluation, image enhancement, and accurate identification of the palm region. For example, image processing techniques (e.g., edge detection, threshold segmentation) can be used to extract and identify the palm region. Furthermore, image resolution evaluation methods (e.g., gradient method, frequency domain analysis) are used to select an image with high resolution and clear palm features, which is then used for subsequent recognition processing.
[0167] After obtaining the recognition image, it is recognized by a local recognition model to obtain a seventh prediction result. Here, the seventh prediction result includes a predicted category and a category score. The predicted category is the category obtained after prediction, and the category score represents the score at which the image is predicted to belong to that category. The higher the category score, the higher the probability that the image will belong to this predicted category. Based on this, if the category score is equal to or greater than the category score threshold (e.g., 0.90), the field terminal determines that the recognition image belongs to the predicted category, and can then perform the corresponding task (e.g., payment processing, access control processing, etc.).
[0168] For example, in one embodiment, the local recognition model can output a category probability distribution corresponding to the image to be recognized. The category corresponding to the highest probability value in the category probability distribution is used as the predicted category of the image to be recognized, and this highest probability value is used as the category score. In another embodiment, the local recognition model can extract feature vectors from the image to be recognized, and then match the extracted feature vectors with feature vectors already existing in the database, and perform feature matching using, for example, the k-nearest neighbors (KNN) algorithm. As a result, the category corresponding to the feature vector with the highest similarity is used as the predicted category of the image to be recognized, and the one with the highest similarity is used as the category. Here, cosine similarity may be used as the similarity between feature vectors.
[0169] Furthermore, embodiments of this application provide a method for performing image recognition locally on a terminal. This method allows the terminal to recognize collected images using a local recognition model. Therefore, it reduces the data processing load on the server side, improves recognition efficiency, and is independent of the network environment.
[0170] In addition to one or more embodiments corresponding to Figure 3 above, another optional embodiment according to the embodiments of this application, after the step of obtaining a seventh prediction result by a local recognition model based on a recognition image, further, If the category score in the seventh prediction result is less than the category score threshold, the recognition image is sent to the server, and the server obtains an image recognition result using an image recognition model based on the recognition image. This may include a step of receiving image recognition results sent from a server.
[0171] In one or more embodiments, a method for performing image recognition on the server side is described. As is clear from the embodiments described above, after the terminal recognizes the image for recognition using a local recognition model, if the obtained category score is less than the category score threshold (e.g., 0.90), the field terminal can send the recognized image to the server.
[0172] Specifically, after receiving an image for recognition, the server calls an image recognition model to recognize the image and obtain the image recognition result. The server then sends the image recognition result to the field terminal. The image recognition result includes the predicted category and category score of the image for recognition. Based on this, the field terminal can perform the corresponding task (e.g., payment processing, access control, etc.) according to the predicted category in the image recognition result.
[0173] For example, in one embodiment, the image recognition model may output a category probability distribution corresponding to the image to be recognized. The category corresponding to the highest probability value in the category probability distribution is used as the predicted category of the image to be recognized, and the highest probability value is used as the category score of the predicted category. In another embodiment, the image recognition model can extract feature vectors from the image to be recognized and then match the extracted feature vectors with feature vectors already existing in the database. As a result, the category corresponding to the feature vector with the highest similarity is used as the predicted category of the image to be recognized, and the category with the highest similarity is used as the category score of the predicted category.
[0174] Furthermore, embodiments of this application provide a method for performing image recognition on the server side. With this method, if the category to which an image belongs cannot be locally predicted on the terminal, the terminal can request the server to predict the image. Since the server-side model has superior recognition capabilities, the success rate and accuracy of image recognition can be improved.
[0175] In conjunction with the above description, the data processing method in this application will now be described from the server's perspective. Referring to Figure 10, the data processing method in the embodiment of this application can be executed by the server alone or by the server in cooperation with at least one terminal. The method of this application includes the following steps.
[0176] Step 310: Receive K images transmitted from the field terminal. Here, K images are taken by the field terminal using the acquisition device under the current field environment, and K is an integer greater than or equal to 1.
[0177] In one or more embodiments, a field terminal calls an image acquisition device to capture multiple images of the current field environment, obtaining K images. The server receives the K images transmitted from the field terminal.
[0178] Step 320: Based on K images, the image recognition model obtains K first prediction results.
[0179] In one or more embodiments, the server inputs each of the K images into an image recognition model, the image recognition model outputs a first prediction result for each image, and thereby obtains K first prediction results.
[0180] Step 320 in this embodiment is the same as step 220 in the embodiment shown in Figure 3 above, so it will not be described in detail here.
[0181] Step 330: K first prediction results are sent to the field terminal, the field terminal constructs a fine-tuning training set according to the K images and the K first prediction results, the model to be trained is fine-tuned on the field terminal using the fine-tuning training set, and in the process of fine-tuning the model to be trained on the field terminal, a second prediction result corresponding to each image is obtained by the model to be trained based on the images included in the fine-tuning training data of each group in the fine-tuning training set, and the model parameters of the model to be trained are updated according to the second prediction result corresponding to each image and the first prediction result of the images in the fine-tuning training set, and the local recognition model and the model tuning parameters corresponding to the local recognition model are obtained. Here, the fine-tuning training set includes K groups of fine-tuning training data, and the fine-tuning training data of each group includes images and the first prediction result of the images.
[0182] In one or more embodiments, the server transmits K first prediction results to the field terminal. Based on this, the field terminal can use the K images and the K first prediction results to train a model.
[0183] Step 330 in this embodiment is the same as steps 230 to 250 in the embodiment shown in Figure 3 above, so it will not be described in detail here.
[0184] Step 340: If the local recognition model meets the model fine-tuning conditions, receive the model adjustment parameters sent from the field terminal.
[0185] In one or more embodiments, if the local recognition model satisfies the model fine-tuning conditions, the field terminal sends model tuning parameters to the server.
[0186] Step 350: If a model tuning parameter set has been obtained from at least one terminal, update the model parameters of the image recognition model. Here, the model tuning parameter set includes the model tuning parameters.
[0187] In one or more embodiments, the server integrates model tuning parameters reported by different terminals to obtain a model tuning parameter set. Based on this, the server uses the model tuning parameter set to update the model parameters of the image recognition model, i.e., fine-tune the image recognition model.
[0188] Steps 340 to 350 in this embodiment are the same as step 260 in the embodiment shown in Figure 3 above, so they will not be described in detail here.
[0189] Embodiments of this application provide a data processing method. This method allows a terminal to fine-tune a local model based on images collected in the field environment and report the fine-tuned model tuning parameters to a server. The server integrates the model tuning parameter sets reported from each terminal to update the image recognition model. This makes the image recognition model applicable to various specific field environments, improving the accuracy of model recognition, saving server processing resources, and improving the efficiency of model learning.
[0190] In addition to one or more embodiments corresponding to Figure 10 above, in another optional embodiment according to the embodiments of this application, if a set of model adjustment parameters is obtained from at least one terminal, the step of updating the model parameters of the image recognition model is specifically: The steps include obtaining a model adjustment parameter set from at least one terminal, The steps include: applying weighting to the model tuning parameter set according to the overall recognition score corresponding to each terminal, and obtaining a weighted model tuning parameter set; This may include a step of updating the model parameters of an image recognition model using a weighted set of model tuning parameters.
[0191] In one or more embodiments, a method for setting the weighting effect based on the voting results of fine-tuning is described. As is evident from the embodiments described above, the server receives model tuning parameters sent from different field terminals and then decides whether to accept the model tuning parameters sent from each terminal based on the overall recognition score corresponding to each terminal. Here, the method for calculating the overall recognition score has been described in the embodiments described above and will not be described in detail here. If the server accepts the model tuning parameters, the model tuning parameters are fused into the image recognition model for training.
[0192] In one feasible form, the server opens an interface to receive model tuning parameters and voting scores (or overall recognition scores) sent from terminals. Upon receiving the voting scores, the server further calculates the overall recognition score. If the overall recognition score is equal to or greater than a score threshold (e.g., 0.5), it means that most terminals voted in favor, and the server accepts these model tuning parameters. If the overall recognition score is lower than the score threshold, it means that most terminals voted against, and the server does not adopt these model tuning parameters. Furthermore, if the server accepts the model tuning parameters, it may also weight them according to the overall recognition score.
[0193] For example, let's assume that "Terminal A" reports one set of model tuning parameters, and the overall recognition score calculated from the voting scores of multiple terminals associated with "Terminal A" is "0.8"; that "Terminal B" reports one set of model tuning parameters, and the overall recognition score calculated from the voting scores of multiple terminals associated with "Terminal B" is "0.3"; and that "Terminal C" reports one set of model tuning parameters, and the overall recognition score calculated from the voting scores of multiple terminals associated with "Terminal C" is "0.9". If the score threshold is 0.5, the model tuning parameters reported by "Terminal B" will not be adopted. The influence of the model tuning parameters reported by "Terminal A" on fine-tuning the image recognition model is "0.8", and this "0.8" is used to weight the model tuning parameters, resulting in weighted model tuning parameters. The influence of the model tuning parameters reported by "Terminal C" on fine-tuning the image recognition model is "0.9". This "0.9" is used to weight the model tuning parameters, and the weighted model tuning parameters are obtained. Thus, the influence of the model tuning parameters reported by "Terminal C" on fine-tuning the image recognition model is even greater.
[0194] Furthermore, embodiments of this application describe a method for setting the weighting influence based on the voting results of fine-tuning. As described above, each terminal has different perceptions of the field environment, so the results of fine-tuning may also differ. Therefore, the server fine-tunes the image recognition model according to the overall recognition score and model tuning parameters reported from terminals in different environments. This allows the server-side image recognition model to incorporate various fine-tuning results, thereby improving and enhancing the model's recognition capabilities.
[0195] In addition to one or more embodiments corresponding to Figure 10 above, in another optional embodiment according to the embodiments of this application, if a set of model adjustment parameters is obtained from at least one terminal, the step of updating the model parameters of the image recognition model is specifically: If a model parameter set is obtained from at least one terminal, update the model parameters of the image recognition model, where the model parameter set includes model tuning parameters, and the model tuning parameters are model parameters. or If a gradient set is obtained from at least one terminal, update the model parameters of the image recognition model, where the gradient set includes model tuning parameters, and the model tuning parameters are gradients. or If an optimization algorithm parameter set is obtained from at least one terminal, the model parameters of the image recognition model may be updated, where the optimization algorithm parameter set includes model tuning parameters, and the model tuning parameters are optimization algorithm parameters.
[0196] In one or more embodiments, three methods for training an image recognition model based on a set of model tuning parameters are presented. As is evident from the embodiments described above, the model tuning parameters include, but are not limited to, model parameters, gradients, and optimization algorithm parameters. Based on this, the server can fine-tune the image recognition model according to the model tuning parameters.
[0197] 1. Based on model parameters The model tuning parameter set may also be a model parameter set. That is, each terminal originating from a different environment can report model parameters to the server. Exemplaryly, the server can fuse these model parameters using a model fusion strategy and then use the fused model parameters to fine-tune the image recognition model.
[0198] Model fusion strategies include, but are not limited to, stacking strategies, boosting strategies, and bootstrap aggregating (bagging).
[0199] 2. Based on gradient The model tuning parameter set may also be a gradient set. That is, each terminal originating from a different environment can report its gradient to the server. Exemplaryly, the server can average these gradients to obtain a gradient mean. Based on this, the server can use the gradient mean to fine-tune the image recognition model.
[0200] 3. Based on optimization algorithm parameters The model tuning parameter set may also be the optimization algorithm parameter set, where the optimization algorithm parameters include, but are not limited to, the optimization algorithm, learning rate, and iteration count. Each terminal originating from a different environment can report its own optimization algorithm parameters to the server. Exemplaryly, the server can integrate these optimization algorithm parameters and fine-tune the image recognition model using the most frequently occurring optimization algorithm parameters.
[0201] It is understandable that the server can select one of the fine-tuning methods reported by the terminal for training. For example, if the fine-tuning method involves changing model parameters, the updated model parameters can be used as is during training. Alternatively, if the fine-tuning method involves changing optimization algorithm parameters, those optimization algorithm parameters can be used when training a large model.
[0202] Furthermore, embodiments of this application provide a method for training an image recognition model based on a set of model tuning parameters. Using this method, each terminal originating from a different environment reports its own model tuning parameters to the server, which can then fine-tune the image recognition model according to the specific details of these parameters. This improves the flexibility and versatility of the model training method.
[0203] In addition to one or more embodiments corresponding to Figure 10 above, in another optional embodiment according to the embodiments of this application, if a model adjustment parameter set is obtained from at least one terminal, after the step of updating the model parameters of the image recognition model, The process may include the steps of sending the model tuning parameters of the image recognition model to at least one terminal, and each terminal within that at least one terminal updating the recognition parameters of the recognition model using the model tuning parameters of the image recognition model.
[0204] In one or more embodiments, a method for updating other recognition models based on an image recognition model is described. As is evident from the embodiments described above, the server can send model tuning parameters of the image recognition model to terminals in different environments so that these terminals can use the model tuning parameters to fine-tune the recognition model. The method for fine-tuning based on the model tuning parameters is described in the embodiments described above, so it is understandable that it will not be described in detail here.
[0205] For ease of understanding, please refer to Figure 11. Figure 11 is a framework diagram for data processing between a terminal and a server in one embodiment of the present application. As shown in the figure, taking a field terminal as an example, the field terminal includes a local recognition module shown as F1, an automatic fine-tuning module based on field environment information shown as F2, a multi-terminal status monitoring module shown as F3, and a field multi-terminal model fine-tuning synchronization module shown as F4. The server includes a field fine-tuning influence module shown as F5 and an image recognition model shown as F6.
[0206] During the model fine-tuning process, the local recognition module can transmit field environment information to the field environment-based automatic fine-tuning module. The field environment-based automatic fine-tuning module then fine-tunes the local model to be trained according to the field environment information and the fine-tuning training set.
[0207] During the status monitoring process, the local recognition module can also transmit the terminal status of the field terminal to the multi-terminal status monitoring module. The multi-terminal status monitoring module then feeds back the stability status of the model to the local recognition module.
[0208] During the synchronization process of fine-tuning results from multiple terminals, the automatic fine-tuning module, based on field environment information, sends model tuning parameters to the field's multi-terminal model fine-tuning synchronization module. This allows other terminals to use these model tuning parameters to fine-tune the local model to be trained and provide feedback on the corresponding fine-tuning results.
[0209] During the server's fine-tuning synchronization process, the field multi-terminal model fine-tuning synchronization module feeds back the voting results from each terminal to the server-side field fine-tuning influence module. The field fine-tuning influence module obtains an overall recognition score from the voting results. If the server decides to accept the model tuning parameters reported from the field terminals according to the overall recognition score, it can fine-tune the image recognition model using the model tuning parameters and then send the resulting model tuning parameters to the field terminals.
[0210] Furthermore, embodiments of this application provide a method for updating other recognition models based on an image recognition model. This method also allows the server to send finely tuned model tuning parameters to each terminal, enabling the terminals to fine-tune their local models based on these parameters. This results in the continuous learning and optimization of the models, improving the overall accuracy and efficiency of image recognition.
[0211] The data processing device described in this application will be described in detail below. Please refer to Figure 12. Figure 12 is a schematic diagram of a data processing device in one embodiment of this application. The data processing device 40 is The current field environment involves a step of acquiring K images using an image acquisition device, wherein the imaging module 410 is configured to perform this step, where K is an integer of 1 or more. A transmission module 420 is configured to send K images to a server, and the server performs the step of obtaining K first prediction results using an image recognition model based on the K images. An acquisition module 430 is configured to perform the steps of constructing a fine-tuning training set according to K images and K first prediction results transmitted from a server, wherein the fine-tuning training set includes K groups of fine-tuning training data, and each group of fine-tuning training data includes images and first prediction results for those images. An update module 440 is configured to perform the following steps: fine-tune the model to be trained on the field terminal using a fine-tuning training set; in the process of fine-tuning the model to be trained on the field terminal, obtain a second prediction result corresponding to each image using the model to be trained based on the images contained in the fine-tuning training data of each group in the fine-tuning training set; update the model parameters of the model to be trained according to the second prediction result corresponding to each image and the first prediction result of the images in the fine-tuning training set; and obtain a local recognition model and model tuning parameters corresponding to the local recognition model. Furthermore, the system includes a transmission module 420 configured to perform the following steps: if the local recognition model satisfies the model fine-tuning conditions, it sends model tuning parameters to a server, and the server updates the model parameters of the image recognition model according to a set of model tuning parameters from at least one terminal, wherein the set of model tuning parameters includes the model tuning parameters.
[0212] Based on the embodiment corresponding to Figure 12 above, in another embodiment of the data processing device 40 according to the embodiment of this application, the data processing device 40 further comprises a receiving module 450, The transmission module 420 is further configured to send a model training request to the server, and the server is configured to perform the step of determining the training dataset to be applied to the field terminal in accordance with the model training request. The receiving module 450 is configured to perform the step of receiving an initial training set sent from the server, wherein the initial training set includes initial training data for M groups, and the initial training data for each group includes images and image labeling results. The acquisition module 430 is further configured to perform the step of acquiring M initial prediction results by the initial recognition model based on M images included in the initial training set, wherein each initial prediction result includes the predicted category and category score of the image. The update module 440 is further configured to perform the step of updating the model parameters of the initial recognition model according to the M initial prediction results and the M labeling results included in the initial training set, thereby obtaining a model to be trained.
[0213] Based on the embodiment corresponding to Figure 12 above, in another embodiment of the data processing device 40 according to the embodiment of this application, the data processing device 40 further comprises a processing module 460, The acquisition module 430 is further configured to perform the step of acquiring field environment information where the field terminal is located, prior to the step of obtaining K images by taking images with the image acquisition device, wherein the field environment information includes at least one of light intensity and background noise. The processing module 460 is configured to perform the following steps in response to a first adjustment operation on the image acquisition device when the light intensity included in the field environment information is not within the light intensity range: adjusting a first application parameter of the image acquisition device, wherein the first application parameter includes at least one of the shutter speed, sensitivity parameter and exposure compensation parameter. The processing module 460 is further configured to perform the step of adjusting a second application parameter of the image acquisition device in response to a second adjustment operation on the image acquisition device, if the background noise included in the field environment information is greater than or equal to a background noise threshold, wherein the second application parameter includes at least one of a sharpness parameter, a sensitivity parameter, and a noise reduction parameter.
[0214] Based on the embodiment corresponding to Figure 12 above, in another embodiment of the data processing device 40 according to the embodiment of this application, the data processing device 40 further comprises a decision module 470, The processing module 460 is further configured to perform the step of obtaining the recognition accuracy of the local recognition model for N images, where N is an integer of 1 or more, and the N images are captured by the image acquisition device. The decision module 470 is further configured to perform a step of determining that the local recognition model satisfies the model fine-tuning conditions if the recognition accuracy is equal to or greater than the accuracy threshold, or the transmission module 420 is further configured to perform a step of sending model tuning parameters to T terminals if the recognition accuracy is equal to or greater than the accuracy threshold, and each of the T terminals updates the model parameters of their respective corresponding models to be trained according to the model tuning parameters, thereby obtaining T recognition models, where the T terminals are associated with field terminals and T is an integer greater than or equal to 1, the acquisition module 430 is further configured to perform a step of acquiring a voting score corresponding to each terminal among the T terminals, the voting score being determined according to the prediction results of the recognition model and the prediction results of the image recognition model, the decision module 470 is further configured to perform a step of determining an overall recognition score according to the voting score corresponding to each terminal, and the decision module 470 is further configured to perform a step of determining that the local recognition model satisfies the model fine-tuning conditions if the overall recognition score is equal to or greater than the recognition score threshold.
[0215] Based on the embodiment corresponding to Figure 12 above, in another embodiment of the data processing device 40 according to the embodiment of this application, The transmission module 420 is further configured to transmit N images captured by the image acquisition device to a server, and the server performs the step of obtaining N third prediction results using an image recognition model based on the N images. The receiving module 450 is further configured to perform the step of receiving N third prediction results sent from the server, The acquisition module 430 is further configured to perform the step of acquiring N fourth prediction results by a local recognition model based on N images. The processing module 460 is further configured to perform the step of verifying N fourth prediction results according to N third prediction results to obtain recognition accuracy for N images.
[0216] Based on the embodiment corresponding to Figure 12 above, in another embodiment of the data processing device 40 according to the embodiment of this application, The decision module 470 specifically performs the steps of aggregating the voting scores of T terminals to obtain a total voting score, The system is configured to perform the steps of obtaining an overall recognition score from the ratio of the total voting score to the T-value.
[0217] Based on the embodiment corresponding to Figure 12 above, in another embodiment of the data processing device 40 according to the embodiment of this application, The decision module 470 specifically takes the step of obtaining a weighting parameter set corresponding to each terminal in T terminals, wherein the weighting parameter set includes at least one of device weighting, environment weighting, and priority weighting. For each of the T terminals, the voting score of the terminal is weighted using the terminal's weighting parameter set to obtain the terminal's weighted voting score. The system is configured to perform the steps of determining an overall recognition score according to the voting weight score corresponding to each terminal in the T terminals.
[0218] Based on the embodiment corresponding to Figure 12 above, in another embodiment of the data processing device 40 according to the embodiment of this application, The decision module 470 is further configured to perform the steps of determining, before the step of sending model adjustment parameters to T terminals, that if the field terminal and at least one terminal are in the same location area, at least one terminal is in an associated relationship with the field terminal, and that at least one terminal is one of the T terminals. or The decision module 470 is further configured to perform the steps of determining that, prior to the step of sending model adjustment parameters to T terminals, if the same binding object is set on the field terminal and at least one terminal, at least one terminal is in an associated relationship with the field terminal, and determining that at least one terminal is one of the T terminals. or The decision module 470 is further configured to perform the steps of determining that, if the field terminal and at least one terminal are connected to the same access point, at least one terminal is associated with the field terminal, and determining that at least one terminal is one of the T terminals, before sending the model adjustment parameters to T terminals.
[0219] Based on the embodiment corresponding to Figure 12 above, in another embodiment of the data processing device 40 according to the embodiment of this application, The acquisition module 430 is further configured to perform the step of acquiring images to be tested after the step of determining that the local recognition model satisfies the model fine-tuning conditions. The acquisition module 430 is further configured to perform a step of obtaining a fifth prediction result by the local recognition model based on the image to be tested; The acquisition module 430 is further configured to perform the step of acquiring T sixth prediction results from T terminals, where each sixth prediction result is acquired from the terminals by the recognition model based on the image to be tested. The processing module 460 is further configured to perform the step of having the local recognition model perform the corresponding task if it is determined that the local recognition model is in a model stable state according to the fifth prediction result and the T sixth prediction results.
[0220] Based on the embodiment corresponding to Figure 12 above, in another embodiment of the data processing device 40 according to the embodiment of this application, The transmitting module 420 is further configured to perform the following steps: after determining an overall recognition score according to the voting score corresponding to each terminal, if the local recognition model does not satisfy the model fine-tuning conditions, it sends a model fine-tuning request to the terminals within T terminals, the terminals update the model parameters of the model to be trained according to the model fine-tuning request, and obtain a recognition model, where the T terminals are related to field terminals and T is an integer greater than or equal to 1. The receiving module 450 is further configured to perform the step of receiving model adjustment parameters sent from the terminal, The update module 440 is further configured to perform the step of updating the model parameters of the model to be trained on the field terminal using the model adjustment parameters sent from the terminal.
[0221] Based on the embodiment corresponding to Figure 12 above, in another embodiment of the data processing device 40 according to the embodiment of this application, The imaging module 410 is further configured to update the model parameters of the model to be trained according to K second prediction results and fine-tuning training sets, and after the step of obtaining a local recognition model and model tuning parameters corresponding to the local recognition model, if the local recognition model satisfies the model tuning conditions, it performs the step of capturing an image for recognition with an image acquisition device. The acquisition module 430 is further configured to perform the step of acquiring a seventh prediction result by a local recognition model based on the recognition image, wherein the seventh prediction result includes a prediction category and a category score. The decision module 470 is further configured to perform the step of determining that the recognition image belongs to a predicted category in the seventh prediction result if the category score in the seventh prediction result is equal to or greater than the category score threshold.
[0222] Based on the embodiment corresponding to Figure 12 above, in another embodiment of the data processing device 40 according to the embodiment of this application, The transmission module 420 is further configured to send the recognition image to the server if, after the step of obtaining a seventh prediction result by the local recognition model based on the recognition image, the category score in the seventh prediction result is less than the category score threshold, and the server then performs the step of obtaining an image recognition result by the image recognition model based on the recognition image. The receiving module 450 is further configured to perform the step of receiving image recognition results sent from the server.
[0223] The data processing device described in this application will be described in detail below. Please refer to Figure 13. Figure 13 is a schematic diagram of a data processing device in one embodiment of this application. The data processing device 50 is A receiving module 510 is configured to perform the step of receiving K images transmitted from a field terminal, wherein the K images are taken by the field terminal with a collection device under the current field environment, and K is an integer of 1 or more. An acquisition module 520 is configured to perform the step of obtaining K first prediction results by an image recognition model based on K images, A transmission module 530 is configured to perform the following steps: transmit K first prediction results to a field terminal; the field terminal constructs a fine-tuning training set according to K images and K first prediction results; fine-tune the model to be trained on the field terminal using the fine-tuning training set; in the process of fine-tuning the model to be trained on the field terminal, obtain a second prediction result corresponding to each image using the model to be trained based on the images included in the fine-tuning training data of each group within the fine-tuning training set; update the model parameters of the model to be trained according to the second prediction result corresponding to each image and the first prediction results of the images in the fine-tuning training set; and obtain a local recognition model and model tuning parameters corresponding to the local recognition model, wherein the fine-tuning training set includes K groups of fine-tuning training data, and each group of fine-tuning training data includes images and first prediction results of the images. Furthermore, if the local recognition model satisfies the model fine-tuning conditions, the receiving module 510 is configured to perform a step of receiving model adjustment parameters transmitted from the field terminal, The system includes an update module 540 configured to perform the step of updating the model parameters of an image recognition model, where the model adjustment parameter set includes model adjustment parameters, if a model adjustment parameter set is obtained from at least one terminal.
[0224] Based on the embodiment corresponding to Figure 13 above, in another embodiment of the data processing device 50 according to the embodiment of this application, The update module 540 specifically includes the steps of obtaining a model adjustment parameter set from at least one terminal, The steps include: applying weighting to the model tuning parameter set according to the overall recognition score corresponding to each terminal, and obtaining a weighted model tuning parameter set; It is configured to perform the steps of updating the model parameters of an image recognition model using a weighted set of model tuning parameters.
[0225] Based on the embodiment corresponding to Figure 13 above, in another embodiment of the data processing device 50 according to the embodiment of this application, The update module 540 specifically performs the step of updating the model parameters of an image recognition model when a model parameter set has been obtained from at least one terminal, wherein the model parameter set includes model tuning parameters, and the model tuning parameters are model parameters. or If a gradient set is obtained from at least one terminal, the step is to update the model parameters of the image recognition model, where the gradient set includes model tuning parameters, and the model tuning parameters are gradients. or If an optimization algorithm parameter set is obtained from at least one terminal, the system is configured to perform a step of updating the model parameters of an image recognition model, where the optimization algorithm parameter set includes model tuning parameters and the model tuning parameters are optimization algorithm parameters.
[0226] Based on the embodiment corresponding to Figure 13 above, in another embodiment of the data processing device 50 according to the embodiment of this application, The transmitting module 530 is further configured such that, if a model adjustment parameter set is obtained from at least one terminal, after the step of updating the model parameters of the image recognition model, it transmits the model adjustment parameters of the image recognition model to at least one terminal, and each terminal within that at least one terminal performs the step of updating the recognition parameters of the recognition model using the model adjustment parameters of the image recognition model.
[0227] Figure 14 is a schematic diagram of a computer device in one embodiment of the present application. The computer device 60 may include an input device 610, an output device 620, a processor 630, and memory 640. In this embodiment, the output device 620 may be a display device.
[0228] Memory 640 includes read-only memory and random-access memory, providing instructions and data to the processor 630. A portion of memory 640 may also include non-volatile random-access memory (NVRAM).
[0229] Memory 640 contains the following elements, executable modules or data structures, or subsets thereof, or extensions thereof: Operation Instructions: Contains operation instructions used to perform an operation. Operating System: Contains system programs used to implement basic business operations and handle hardware-based tasks.
[0230] The processor 630 controls and performs operations on the computer device 60. Here, the processor 630 is also called the central processing unit (CPU).
[0231] In a specific application, the various components of the computer device 60 are interconnected via a bus system 650. Here, the bus system 650 may include, in addition to the data bus, a power bus, a control bus, a status signal bus, and so on. For ease of explanation, in the drawings, the various buses are labeled as the bus system 650.
[0232] The method according to the embodiments of the present application may be applied to the processor 630 or may be implemented by the processor 630. The processor 630 may be an integrated circuit chip with signal processing functions. During implementation, each step of the above method can be completed by an integrated logic circuit in hardware form or an instruction in software form within the processor 630. The processor 630 is a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, or discrete gate hardware components, and can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or other conventional processors. The steps of the method disclosed in the embodiments of the present application may be directly implemented by the hardware in the processor, or may be implemented by a combination of hardware modules and software modules in the processor. The software module is disposed in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. The storage medium is disposed in the memory 640, and the processor 630 reads the information in the memory 640 and combines it with its hardware to complete the steps of the above method.
[0233] In the above embodiments, the steps executed by the terminal or the server may be based on the structure of the computer device shown in FIG. 14.
[0234] In an embodiment of the present application, there is also provided a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the steps of the methods described in each of the above embodiments are realized.
[0235] In an embodiment of the present application, there is also provided a computer program product including a computer program, wherein when the computer program is executed by a processor, the steps of the methods described in each of the above embodiments are realized.
[0236] It should be understood that in a specific embodiment of the present application, related data such as palm print information, iris information, and face information are included. When applying the above embodiments of the present application to a specific product or technology, user permission or consent is required, and the collection, use, and processing of related data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0237] A person skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific operation processes of the systems, apparatuses, and units described above may refer to the corresponding processes in the method embodiments described above, and thus will not be described in detail here.
[0238] In some embodiments of this application, it should be understood that the systems, devices, and methods disclosed above may be implemented in other ways. For example, the embodiments of the devices described above are merely examples. For example, the division of the units is merely a division of logical functions, and there may be other methods of division in actual implementations. For example, multiple units or components may be coupled or integrated into another system, or some functions may be ignored or not performed. Also, the mutual coupling, direct coupling, or communication connection illustrated or described may be an indirect coupling or communication connection via some interface, device, or unit, which may be electrical, mechanical, or otherwise.
[0239] The units described as individual components may or may not be physically separated, and the components shown as units may or may not be physical units. In other words, they may be located in one place or distributed across multiple network units. To achieve the objectives of the technical method of this embodiment, some or all of the units can be selected according to the actual needs.
[0240] Furthermore, each functional unit in each embodiment of this application may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The above-mentioned integrated unit can be implemented in the form of hardware or a software functional unit.
[0241] When the aforementioned integrated unit is implemented in the form of a software function unit and sold or used as an independent product, it can be stored on a computer-readable storage medium. Based on this understanding, any part of the technical method of this application that can be said to be essentially or contribute to the prior art, or all or part of the technical method, can be embodied in the form of a software product. A computer software product is stored on a storage medium and includes a number of instructions that enable a computer device (such as a server or terminal device) to perform all or part of the steps of the method described in each embodiment of this application. The aforementioned storage mediums include various media capable of storing computer programs, such as U disks, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0242] As described above, the embodiments described above are used not to limit the technical methods of this application, but solely to illustrate the technical methods of this application. Although this application has been described in detail with reference to the embodiments described above, those skilled in the art will understand that the technical methods described in the embodiments above can be modified or some of the technical features therein can be replaced with equivalent ones. However, these modifications or replacements do not cause the essence of the corresponding technical methods to deviate from the spirit and scope of the technical methods described in the embodiments of this application.
Claims
1. A data processing method performed by a field terminal, The current field environment involves the steps of obtaining K (where K is an integer greater than or equal to 1) images by taking pictures with an image acquisition device, The steps include sending the K images to a server, and the server obtaining K first prediction results based on the K images using an image recognition model, A step of constructing a fine-tuning training set according to the K images and the K first prediction results transmitted from the server, wherein the fine-tuning training set includes K groups of fine-tuning training data, and each group of fine-tuning training data includes an image and a first prediction result of the image. The steps include: fine-tuning the model to be trained at the field terminal using the fine-tuning training set; in the process of fine-tuning the model to be trained at the field terminal, obtaining a second prediction result corresponding to each image using the model to be trained based on the images included in the fine-tuning training data of each group within the fine-tuning training set; updating the model parameters of the model to be trained according to the second prediction result corresponding to each image and the first prediction result of the images within the fine-tuning training set, thereby obtaining a local recognition model and model adjustment parameters corresponding to the local recognition model; A data processing method comprising the steps of: sending the model adjustment parameters to a server if the local recognition model satisfies model fine-tuning conditions, and the server updating the model parameters of the image recognition model according to a set of model adjustment parameters from at least one terminal, wherein the set of model adjustment parameters includes the model adjustment parameters.
2. The steps include sending a model training request to a server, and the server determining a training dataset to be applied to the field terminal in accordance with the model training request, A step of receiving an initial training set transmitted from the server, wherein the initial training set includes initial training data for M groups, and the initial training data for each group includes images and the labeling results of the images. A step of obtaining M initial prediction results by an initial recognition model based on M images included in the initial training set, wherein each initial prediction result includes the predicted category of the image and a category score. The data processing method according to claim 1, further comprising the steps of updating the model parameters of the initial recognition model according to the M initial prediction results and the M labeling results included in the initial training set, to obtain the model to be trained.
3. In the current field environment, before the step of obtaining K images by taking pictures with an image acquisition device, A step of acquiring on-site environmental information where the aforementioned on-site terminal is located, wherein the on-site environmental information includes at least one of light intensity and background noise. If the light intensity included in the aforementioned field environment information is not within the light intensity range, the step of adjusting a first application parameter of the image acquisition device in response to a first adjustment operation on the image acquisition device, wherein the first application parameter includes at least one of shutter speed, sensitivity parameter, and exposure compensation parameter. The data processing method according to claim 1, further comprising the step of adjusting a second application parameter of the image acquisition device in response to a second adjustment operation on the image acquisition device when the background noise included in the field environment information is greater than or equal to a background noise threshold, wherein the second application parameter includes at least one of a sharpness parameter, a sensitivity parameter, and a noise reduction parameter.
4. After the step of updating the model parameters of the model to be trained according to the second prediction result corresponding to each of the aforementioned images and the first prediction result of the images in the fine-tuning training set, and obtaining the local recognition model and the model tuning parameters corresponding to the local recognition model, A step of obtaining the recognition accuracy of the local recognition model for N (where N is an integer of 1 or more) images, wherein the N images were captured by the image acquisition device, If the recognition accuracy is equal to or greater than the accuracy threshold, the local recognition model is determined to satisfy the model fine-tuning conditions. Alternatively, the data processing method according to claim 1, further comprising the steps of: if the recognition accuracy is equal to or greater than an accuracy threshold, transmitting the model adjustment parameters to T (T is an integer of 1 or more) terminals associated with the field terminal, and each of the T terminals updating the model parameters of their respective corresponding models to be trained according to the model adjustment parameters, thereby obtaining T recognition models; obtaining a voting score corresponding to each terminal among the T terminals, wherein the voting score is determined according to the prediction results of the recognition model and the prediction results of the image recognition model; determining an overall recognition score according to the voting score corresponding to each terminal; and determining that the local recognition model satisfies the model fine-tuning conditions if the overall recognition score is equal to or greater than a recognition score threshold.
5. The step of obtaining the recognition accuracy of the local recognition model for N images is: The steps include: transmitting the N images captured by the image acquisition device to the server, and the server obtaining N third prediction results based on the N images using the image recognition model; The steps include receiving the N third prediction results transmitted from the server, The steps include obtaining N fourth prediction results using the local recognition model based on the N images, The data processing method according to claim 4, comprising the step of verifying the N fourth prediction results according to the N third prediction results and obtaining the recognition accuracy of the N images.
6. The step of determining an overall recognition score according to the voting score corresponding to each of the aforementioned terminals is as follows: The steps include: aggregating the voting scores of the T terminals to obtain a total voting score; The data processing method according to claim 4, comprising the step of obtaining the overall recognition score from the ratio of the total voting score and the T-value.
7. The step of determining an overall recognition score according to the voting score corresponding to each of the aforementioned terminals is as follows: A step of obtaining a weighting parameter set corresponding to each terminal among the T terminals, wherein the weighting parameter set includes at least one of device weighting, environment weighting, and priority weighting. For each of the T terminals, the voting score of the terminal is weighted using the terminal's weighting parameter set to obtain the weighted voting score of the terminal. The data processing method according to claim 4, comprising the step of determining the overall recognition score according to the voting weight score corresponding to each terminal in the T terminals.
8. Before the step of transmitting the aforementioned model adjustment parameters to T terminals, If the aforementioned field terminal and at least one terminal are located in the same area, the step of determining that the at least one terminal is associated with the field terminal and determining that at least one terminal is one of the T terminals, or If the same binding object is set on the aforementioned field terminal and at least one terminal, the step of determining that the at least one terminal is associated with the aforementioned field terminal and determining that at least one terminal is one of the T terminals, or The data processing method according to claim 4, further comprising the steps of determining that, if the field terminal and at least one terminal are connected to the same access point, the at least one terminal is associated with the field terminal, and determining that the at least one terminal is one of the T terminals.
9. After determining that the local recognition model satisfies the model fine-tuning conditions, Steps to obtain the image to be tested, The steps include obtaining a fifth prediction result using the local recognition model based on the image to be tested, A step of obtaining T sixth prediction results from the T terminals, wherein each sixth prediction result is obtained by the recognition model based on the image to be tested at each terminal. The data processing method according to claim 4, further comprising the step of performing a corresponding task using the local recognition model if it is determined that the local recognition model is in a model-stable state according to the fifth prediction result and the T sixth prediction results.
10. If the local recognition model does not satisfy the model fine-tuning conditions, a model fine-tuning request is sent to T (T is an integer of 1 or more) terminals associated with the field terminal, and the terminals update the model parameters of the model to be trained according to the model fine-tuning request, thereby obtaining a recognition model. The steps include receiving model adjustment parameters transmitted from the aforementioned terminal, The data processing method according to claim 1, further comprising the step of updating the model parameters of the model to be trained at the field terminal using the model adjustment parameters transmitted from the terminal.
11. After the step of updating the model parameters of the model to be trained according to the K second prediction results and the fine-tuning training set, and obtaining the local recognition model and the model tuning parameters corresponding to the local recognition model, If the local recognition model satisfies the model fine-tuning conditions, the step is to capture an image with the image acquisition device to obtain a recognition image, A step of obtaining a seventh prediction result by the local recognition model based on the recognition image, wherein the seventh prediction result includes a prediction category and a category score. The data processing method according to any one of claims 1 to 10, further comprising the step of determining that the recognition image belongs to a predicted category in the seventh prediction result if the category score in the seventh prediction result is equal to or greater than a category score threshold.
12. After the step of obtaining a seventh prediction result using the local recognition model based on the recognition image, If the category score in the prediction result of the seventh step is less than the category score threshold, the recognition image is sent to the server, and the server obtains an image recognition result using the image recognition model based on the recognition image. The data processing method according to claim 11, further comprising the step of receiving the image recognition result transmitted from the server.
13. A data processing method performed by a server, The steps include receiving K (where K is an integer of 1 or more) images transmitted from a field terminal, wherein the K images were taken by the field terminal with a collection device under the current field environment, The steps include obtaining K first prediction results using an image recognition model based on the aforementioned K images, The steps include: transmitting the K first prediction results to the field terminal; the field terminal constructing a fine-tuning training set according to the K images and the K first prediction results; fine-tuning the model to be trained at the field terminal using the fine-tuning training set; and in the process of fine-tuning the model to be trained at the field terminal, obtaining a second prediction result corresponding to each image by the model to be trained based on the images included in the fine-tuning training data of each group within the fine-tuning training set; updating the model parameters of the model to be trained according to the second prediction result corresponding to each image and the first prediction result of the images in the fine-tuning training set to obtain a local recognition model and model adjustment parameters corresponding to the local recognition model, wherein the fine-tuning training set includes K groups of fine-tuning training data, and each group of fine-tuning training data includes images and the first prediction result of the images. If the local recognition model satisfies the model fine-tuning conditions, the step is to receive the model adjustment parameters transmitted from the field terminal. A data processing method comprising: updating the model parameters of an image recognition model if a model adjustment parameter set is obtained from at least one terminal, wherein the model adjustment parameter set includes the model adjustment parameters.
14. If a model adjustment parameter set is obtained from at least one terminal, the step of updating the model parameters of the image recognition model is: The steps include obtaining a model adjustment parameter set from at least one of the terminals, The steps include: applying weighting to the model tuning parameter set according to the overall recognition score corresponding to each terminal, and obtaining a weighted model tuning parameter set; The data processing method according to claim 13, comprising the step of updating the model parameters of the image recognition model using the weighted set of model adjustment parameters.
15. If a model adjustment parameter set is obtained from at least one terminal, the step of updating the model parameters of the image recognition model is: If a model parameter set is obtained from at least one terminal, the step of updating the model parameters of the image recognition model, wherein the model parameter set includes the model adjustment parameters, and the model adjustment parameters are model parameters. or If a gradient set is obtained from at least one terminal, the step of updating the model parameters of the image recognition model, wherein the gradient set includes the model adjustment parameters, and the model adjustment parameters are gradients. or The data processing method according to claim 13, comprising the step of updating the model parameters of the image recognition model when an optimization algorithm parameter set has been obtained from at least one terminal, wherein the optimization algorithm parameter set includes the model adjustment parameters and the model adjustment parameters are optimization algorithm parameters.
16. If a model adjustment parameter set is obtained from at least one terminal, after the step of updating the model parameters of the image recognition model, A data processing method according to any one of claims 13 to 15, comprising the steps of transmitting the model adjustment parameters of the image recognition model to at least one terminal, and each terminal within the at least one terminal updating the recognition parameters of the recognition model using the model adjustment parameters of the image recognition model.
17. A data processing device installed at a field terminal, A shooting module configured to perform the step of capturing images with an image acquisition device and obtaining K (where K is an integer of 1 or more) images in the current field environment, A transmission module is configured to send the K images to a server, and for the server to perform the step of obtaining K first prediction results using an image recognition model based on the K images. An acquisition module is configured to perform the steps of constructing a fine-tuning training set according to the K images and the K first prediction results transmitted from the server, wherein the fine-tuning training set includes K groups of fine-tuning training data, and each group of fine-tuning training data includes an image and a first prediction result for the image. The update module is configured to perform the steps of: fine-tuning the model to be trained at the field terminal using the fine-tuning training set; in the process of fine-tuning the model to be trained at the field terminal, obtaining a second prediction result corresponding to each image using the model to be trained based on the images included in the fine-tuning training data of each group in the fine-tuning training set; updating the model parameters of the model to be trained according to the second prediction result corresponding to each image and the first prediction result of the images in the fine-tuning training set; and obtaining a local recognition model and model adjustment parameters corresponding to the local recognition model. The transmission module is further configured to transmit the model tuning parameters to a server if the local recognition model satisfies model tuning conditions, and the server updates the model parameters of the image recognition model according to a set of model tuning parameters from at least one terminal, wherein the set of model tuning parameters includes the model tuning parameters.
18. A data processing device located on a server, A receiving module is configured to perform the step of receiving K (K is an integer of 1 or more) images transmitted from a field terminal, wherein the K images were taken by the field terminal with a collection device under the current field environment. An acquisition module is configured to perform the step of obtaining K first prediction results using an image recognition model based on the aforementioned K images. A transmission module is configured to perform the following steps: transmit the K first prediction results to the field terminal; the field terminal constructs a fine-tuning training set according to the K images and the K first prediction results; fine-tune the model to be trained at the field terminal using the fine-tuning training set; and in the process of fine-tuning the model to be trained at the field terminal, obtain a second prediction result corresponding to each image using the model to be trained based on the images included in the fine-tuning training data of each group in the fine-tuning training set; update the model parameters of the model to be trained according to the second prediction result corresponding to each image and the first prediction result of the images in the fine-tuning training set to obtain a local recognition model and model tuning parameters corresponding to the local recognition model, wherein the fine-tuning training set includes K groups of fine-tuning training data, and each group of fine-tuning training data includes images and the first prediction result of the images. The process includes a step of updating the model parameters of the image recognition model if a model adjustment parameter set is obtained from at least one terminal, wherein the update module is configured to perform a step of updating the model parameters of the image recognition model, wherein the model adjustment parameter set includes the model adjustment parameters. The receiving module is further configured to perform the step of receiving the model adjustment parameters transmitted from the field terminal if the local recognition model satisfies the model fine-tuning conditions, and is a data processing device.
19. A computer device including memory in which computer programs are stored and a processor, A computer device that, when the computer program is executed, causes the processor to implement the steps of the data processing method described in any one of claims 1 to 12, or to implement the steps of the data processing method described in any one of claims 13 to 16.
20. A computer-readable storage medium on which computer programs are stored, A computer-readable storage medium that, when the computer program is executed, causes the processor to implement a step according to any one of claims 1 to 12, or a step according to any one of claims 13 to 16.
21. A computer program that, when executed by a processor, causes the steps of the data processing method described in any one of claims 1 to 12, or causes the steps of the data processing method described in any one of claims 13 to 16, to be realized.