Image processing method and system
By working collaboratively between the edge and the cloud, and using video datasets to train models, the problem of low model update efficiency was solved, enabling timely model updates and improved efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA (CHINA) CO LTD
- Filing Date
- 2022-08-12
- Publication Date
- 2026-05-15
AI Technical Summary
In existing technologies, the large amount of data leads to low model update efficiency, making it impossible to update the model in a timely manner and affecting the efficiency of image processing.
By working collaboratively between the edge and the cloud, the original cloud model is trained using video datasets obtained from the edge, and the target edge model is trained using video datasets obtained from the cloud, thus achieving collaborative model updates.
Without the need for manual annotation, timely model updates were achieved by coordinating the updates of the original cloud model and edge model using a large amount of video data, thus improving update efficiency.
Smart Images

Figure CN115346104B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent transportation, and more specifically, to an image processing method and system. Background Technology
[0002] Currently, monitoring data is often stored on the client's intranet. Typically, the data is labeled before the model is trained. However, due to the large amount of data, the data processing time is too long, and the model cannot be updated in a timely manner, resulting in the technical problem of low efficiency in updating the model.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This invention provides an image processing method and system to at least address the technical problem of low efficiency in model updates.
[0005] According to one aspect of the present invention, an image processing method is provided, comprising: acquiring an original cloud model from an edge device, wherein the original cloud model is obtained by training an original edge model with a video dataset, the video dataset being obtained by the edge device through monitoring with different image acquisition devices at a first moment; performing image recognition on an input image in a scene to be monitored based on the original cloud model to obtain a feature vector; training a target cloud model based on the feature vector; and training the target cloud model with a video dataset monitored by different image acquisition devices at a second moment to obtain a target edge model, wherein the target edge model is used to update the original edge model, so that the edge device can recognize the video dataset monitored by different image acquisition devices at a third moment based on the target edge model.
[0006] According to one aspect of the present invention, another image processing method is provided, comprising: acquiring a video dataset monitored by different image acquisition devices at a first moment; training an original edge model using the monitored video dataset to obtain an original cloud model, wherein the original cloud model is used to enable the cloud to train feature vectors of an image based on an input image in the scene to be monitored, thereby obtaining a target cloud model; acquiring a target edge model obtained by training the target cloud model using video datasets monitored by different image acquisition devices at a second moment; updating the original edge model to the target edge model, wherein the target edge model is used to identify the video datasets monitored by different image acquisition devices at a third moment.
[0007] According to one aspect of the present invention, an image processing system is provided, comprising: an edge, configured to acquire video datasets monitored by different image acquisition devices at a first moment, and to train an original edge model using the monitored video datasets to obtain an original cloud model; a cloud, configured to perform image recognition on an input image in a scene to be monitored based on the original cloud model to obtain feature vectors, to train a target cloud model based on the feature vectors, and to train the target cloud model using video datasets monitored by different image acquisition devices at a second moment to obtain a target edge model, wherein the target edge model is used to update the original edge model; the edge is further configured to identify video datasets monitored by different image acquisition devices at a third moment based on the target edge model.
[0008] According to one aspect of the present invention, another image processing method is provided, comprising: acquiring an original cloud model from an edge device, wherein the original cloud model is obtained by training an original edge model with a video dataset, the video dataset being obtained by the edge device monitoring traffic roads at a first moment using different image acquisition devices, and the video dataset containing at least one vehicle traveling through the traffic road; identifying input images in the traffic road based on the original cloud model to obtain feature vectors; training a target cloud model based on the feature vectors; and training the target cloud model with video datasets monitored by different image acquisition devices at a second moment to obtain a target edge model, wherein the target edge model is used to update the original edge model, such that the edge device identifies video datasets monitored by different image acquisition devices on the traffic road at a third moment based on the target edge model.
[0009] According to one aspect of the present invention, another image processing method is provided, comprising: displaying an input image of a scene to be monitored on the display screen of a virtual reality (VR) device or an augmented reality (AR) device; the VR device or AR device sending the input image to an edge device, wherein the edge device's original cloud model performs image recognition on the input image and trains a target cloud model based on the recognized feature vectors, the original cloud model being obtained by training the original edge device model with a video dataset, the video dataset being obtained by the edge device using different image acquisition devices at a first moment; after training the target cloud model with the video dataset monitored by different image acquisition devices at a second moment, and updating the original edge device model based on the trained target edge device model, driving the VR device or AR device to render and display the video dataset monitored at a third moment.
[0010] According to one aspect of the present invention, an image processing apparatus is provided, comprising: a first acquisition unit, configured to acquire an original cloud model from an edge device, wherein the original cloud model is obtained by training an original edge model with a video dataset, the video dataset being obtained by the edge device through monitoring with different image acquisition devices at a first moment; a first recognition unit, configured to perform image recognition on an input image in a scene to be monitored based on the original cloud model, thereby obtaining a feature vector; a first training unit, configured to train a target cloud model based on the feature vector; and a second training unit, configured to train the target cloud model with video datasets monitored by different image acquisition devices at a second moment, thereby obtaining a target edge model, wherein the target edge model is used to update the original edge model, such that the edge device can recognize video datasets monitored by different image acquisition devices at a third moment based on the target edge model.
[0011] According to one aspect of the present invention, another image processing apparatus is provided, comprising: a second acquisition unit for acquiring video datasets monitored by different image acquisition devices at a first moment; a third training unit for training an original edge model using the monitored video datasets to obtain an original cloud model, wherein the original cloud model is used to enable the cloud to train feature vectors of images based on input images in the scene to be monitored, thereby obtaining a target cloud model; a third acquisition unit for acquiring a target edge model obtained by training the target cloud model using video datasets monitored by different image acquisition devices at a second moment; and an update unit for updating the original edge model to the target edge model, wherein the target edge model is used to identify video datasets monitored by different image acquisition devices at a third moment.
[0012] According to one aspect of the present invention, another image processing apparatus is provided, comprising: a third acquisition unit, configured to acquire an original cloud model from an edge device, wherein the original cloud model is obtained by training an original edge model with a video dataset, the video dataset being obtained by the edge device monitoring traffic roads at a first moment using different image acquisition devices, and the video dataset containing at least one vehicle traveling through the traffic road; a second recognition unit, configured to recognize an input image in the traffic road based on the original cloud model to obtain a feature vector; a third training unit, configured to train a target cloud model based on the feature vector; and a fourth training unit, configured to train the target cloud model with the video dataset monitored by different image acquisition devices at a second moment to obtain a target edge model, wherein the target edge model is used to update the original edge model, such that the edge device recognizes the video datasets monitored by different image acquisition devices on the traffic road at a third moment based on the target edge model.
[0013] According to one aspect of the present invention, another image processing apparatus is provided, comprising: a presentation unit for displaying an input image in a scene to be monitored on a presentation screen of a virtual reality (VR) device or an augmented reality (AR) device; a third recognition unit for sending the input image to an edge device via the VR device or AR device, wherein the edge device's original cloud model performs image recognition on the input image and trains a target cloud model based on the recognized feature vectors, the original cloud model being obtained by training the original edge model with a video dataset, the video dataset being obtained by the edge device using different image acquisition devices for monitoring at a first moment; and a fifth training unit for training the target cloud model with the video dataset monitored at a second moment using different image acquisition devices, updating the original edge model based on the trained target edge model, and then driving the VR device or AR device to render and display the video dataset monitored at a third moment.
[0014] According to another aspect of the present invention, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored program, wherein, when the program is executed, it controls the device where the storage medium is located to perform the image processing method described above.
[0015] According to another aspect of the present invention, a processor is also provided, which is used to run a program, wherein the image processing method of any one of the above-mentioned methods is executed when the program is running.
[0016] In this embodiment of the invention, an original cloud model from the edge is obtained. This original cloud model is trained using a video dataset, which is obtained by the edge through monitoring with different image acquisition devices at a first-moment interval. Image recognition is performed on the input image in the scene to be monitored based on the original cloud model to obtain feature vectors. A target cloud model is then trained based on these feature vectors. Finally, the target cloud model is trained using video datasets monitored by different image acquisition devices at a second-moment interval to obtain a target edge model. This target edge model is used to update the original edge model, enabling the edge to recognize video datasets monitored by different image acquisition devices at a third-moment interval based on the target edge model. In other words, this embodiment of the invention, without requiring manual annotation, obtains the original cloud model by training the original edge model using the video dataset monitored at the first-moment interval, and obtains the target edge model by training the target cloud model using the video dataset monitored at the second-moment interval. By coordinating the updates of the original cloud model and the original edge model using a large amount of video data, the technical effect of timely model updates is achieved, solving the technical problem of low model update efficiency. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0018] Figure 1(a) is a schematic diagram of an image processing system according to an embodiment of the present invention;
[0019] Figure 1(b) is a schematic diagram of a cloud-edge co-evolution system according to an embodiment of the present invention;
[0020] Figure 2 This is a flowchart of an image processing method according to an embodiment of the present invention;
[0021] Figure 3 This is a flowchart of another image processing method according to an embodiment of the present invention;
[0022] Figure 4 This is a flowchart of an image processing method according to an embodiment of the present invention;
[0023] Figure 5 This is a schematic diagram of the hardware environment of a virtual reality device according to an embodiment of the present invention, which describes an image processing method.
[0024] Figure 6 This is a flowchart of another image processing method according to an embodiment of the present invention;
[0025] Figure 7 This is a schematic diagram of another image processing result according to an embodiment of the present invention;
[0026] Figure 8 This is a schematic diagram of a contrastive learning method according to an embodiment of the present invention;
[0027] Figure 9 This is a schematic diagram of a self-supervised training framework based on a language image pre-training paradigm according to an embodiment of the present invention;
[0028] Figure 10 This is a flowchart of a self-supervised training method for edge models based on a language image pre-training paradigm according to an embodiment of the present invention;
[0029] Figure 11 This is a schematic diagram of a cloud-based multi-model distillation module according to an embodiment of the present invention;
[0030] Figure 12 This is a flowchart of a cloud-based multi-model distillation method according to an embodiment of the present invention;
[0031] Figure 13 This is a schematic diagram of an edge model iterative optimization method according to an embodiment of the present invention;
[0032] Figure 14 This is a structural block diagram of a computing environment according to an embodiment of the present invention;
[0033] Figure 15 This is a structural block diagram of a service grid for an image processing method according to an embodiment of the present invention;
[0034] Figure 16 This is a schematic diagram of an image processing apparatus according to an embodiment of the present invention;
[0035] Figure 17 This is a schematic diagram of an image processing apparatus according to an embodiment of the present invention;
[0036] Figure 18 This is a schematic diagram of another image processing apparatus according to an embodiment of the present invention;
[0037] Figure 19 This is a schematic diagram of another image processing apparatus according to an embodiment of the present invention;
[0038] Figure 20 This is a structural block diagram of a computer terminal according to an embodiment of the present invention. Detailed Implementation
[0039] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0040] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0041] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:
[0042] Self-supervised learning can train networks on unlabeled data by leveraging the characteristics of the data itself. In self-supervised learning, the supervision signal comes from the content of the data itself, that is, it can be a supervision signal that the user gives themselves.
[0043] Supervised learning allows for network training on labeled data by leveraging the inherent characteristics of the data itself.
[0044] Contrastive Language-Image Pre-Training (CLIP) can be a multimodal self-supervised data training method.
[0045] Streaming data, which can be distinguished from traditional static datasets, can be sequential, continuous, and quickly accessible data sequences.
[0046] RocketMQ is a distributed message middleware that can be applied to scenarios such as asynchronous decoupling and peak shaving.
[0047] The open-source stream processing platform (Kafka) can serve as a production-grade consumer middleware, providing a unified, high-throughput, and low-latency platform for processing real-time data.
[0048] Relative entropy (Kullback-Leibler divergence, or simply KL divergence) can be the KL loss function. It is an asymmetric measure of the difference between two probability distributions and can also be called information divergence.
[0049] Transfer learning is a learning method that allows data from an existing model and target domain to be transferred and adapted to data in the target domain.
[0050] Pseudo-labels: Labels obtained by inferring from unlabeled data using existing network models.
[0051] Example 1
[0052] According to an embodiment of the present invention, an embodiment of an image processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0053] This application provides an image processing system as shown in FIG1(a) in its embodiments. It should be noted that the image processing system of this embodiment can be executed by the mobile terminal of the embodiment shown in FIG1(a). FIG1(a) is a schematic diagram of an image processing system according to an embodiment of the present invention. As shown in FIG1(a), the image processing system 1000 may include: edge 1001 and cloud 1002.
[0054] Edge 1001 is used to acquire video datasets monitored by different image acquisition devices at the first moment, and to train the original edge model using the monitored video datasets to obtain the original cloud model.
[0055] Cloud 1001 is used to identify input images in the scene to be monitored based on the original cloud model to obtain feature vectors. The target cloud model is then trained based on the feature vectors. Furthermore, the target cloud model is trained using video datasets monitored by different image acquisition devices at the second time step to obtain the target edge model. The target edge model is used to update the original edge model.
[0056] Edge1001 is also used to identify video datasets monitored by different image acquisition devices at the third time based on the target edge model.
[0057] Optionally, Figure 1(b) is a schematic diagram of a cloud-edge co-evolutionary system according to an embodiment of the present invention. As shown in Figure 1(b), the system may include an edge model self-supervised training module 103 in the edge and a cloud multi-model distillation module 104 in the cloud. The edge model self-supervised training module 103 can construct a streaming data loader, enabling the edge small model to continuously acquire streaming data. Based on the acquired streaming data, the edge small model is self-supervised pre-trained to obtain the final cloud large model. The cloud multi-model distillation module 104 obtains the aforementioned cloud large model through model backflow. The model can process input images under different scene models to obtain feature vectors. It performs average pooling on all feature vectors under different scene models and fuses them to obtain feature vectors representing the image. The fused feature vectors are used as supervision information to conduct supervised training on a large cloud model to obtain a large cloud model+ (which can be the target cloud model). Then, by performing transfer learning on the large cloud model, an edge small model+ (which can be the target edge model) is obtained. The edge small model+ is used to update the edge small model, and this process is repeated to achieve collaborative updating of cloud and edge models, thereby improving the model update efficiency.
[0058] In the image processing system provided in this embodiment of the invention, a video dataset monitored at a first moment is acquired at the edge, and the original edge model is trained using the monitored video dataset to obtain an original cloud model. The input image in the original cloud model is identified through the cloud to obtain feature vectors, and a target cloud model is trained based on the feature vectors. At the same time, the target cloud model is trained using a video dataset monitored at a second moment to obtain a target edge model. Thus, by using a large amount of video to coordinate and update the original cloud model and the original edge model without manual annotation, the technical effect of timely model updating is achieved, solving the technical problem of low efficiency in model updating.
[0059] In the operating environment shown in Figure 1, this embodiment of the invention provides a method from the cloud side as follows: Figure 2 The image processing method shown. Figure 2 This is a flowchart of an image processing method according to an embodiment of the present invention, such as... Figure 2 As shown, the method may include the following steps:
[0060] Step S202: Obtain the original cloud model from the edge, wherein the original cloud model is obtained by training the original edge model with the video dataset, and the video dataset is obtained by monitoring the edge with different image acquisition devices at the first moment.
[0061] In the technical solution provided in step S202 of the present invention, the cloud obtains the original cloud model from the edge. This original cloud model can be obtained by training the original edge model on the edge based on the video dataset. The original edge model can be obtained by the edge from the customer's online environment using the aforementioned video dataset. Optionally, the video dataset in this embodiment can be streaming data obtained from the edge, which may include video data collected by different image acquisition devices in various scenarios, such as video data from urban scenes or rural scenes, and may include multiple frames of images to be monitored; no specific limitation is made here. The image acquisition device can be a device used to acquire images, such as a camera or camcorder; no specific limitation is made here. The original edge model can be used to represent a small edge model, and the original cloud model can be a large cloud model used to represent the data distribution.
[0062] Optionally, the aforementioned video dataset is monitored at the edge through different image acquisition devices in the first moment. This video dataset is used to train the original edge model to obtain the original cloud model, and the cloud obtains the original cloud model that flows back from the edge.
[0063] Step S204: Based on the original cloud model, perform image recognition on the input image in the scene to be monitored to obtain the feature vector.
[0064] In the technical solution provided by step S204 of the present invention, an input image of the scene to be monitored is obtained, and the input image is image recognized by the original cloud model to obtain a feature vector. The scene to be monitored can be a scene collected by an image acquisition device, or a scene to be monitored in advance, such as a city scene, a railway scene, etc. This is only an example and no specific limitation is made. The input image can be each frame of the image to be detected in the video dataset.
[0065] For example, the scene to be monitored is an urban scene of scene 1, scene 2, scene 3... scene N. The video datasets of scene 1, scene 2, scene 3... scene N are obtained, the video datasets are processed to obtain the input images of the scene to be monitored, and the original cloud model is used to identify the input images in the scene to be monitored to obtain the feature vector.
[0066] Step S206: The target cloud model is obtained by training based on the feature vector.
[0067] In the technical solution provided by step S206 of the present invention, the feature vector is obtained, and the original cloud model can be trained based on the feature vector to obtain the target cloud model, wherein the target cloud model can be a large cloud model.
[0068] In this embodiment of the invention, a multi-model distillation module can be used to distill the feature vectors obtained from the original cloud model for multiple scenarios to train a target cloud model. Since the data in multiple scenarios are fitted during the model training process, the accuracy of the model in processing data in different scenarios is improved, thereby improving the generalization ability of the target cloud model.
[0069] Step S208: The target cloud model is trained using video datasets monitored by different image acquisition devices at the second time to obtain the target edge model. The target edge model is used to update the original edge model, so that the edge can identify the video datasets monitored by different image acquisition devices at the third time based on the target edge model.
[0070] In the technical solution provided by step S208 of the present invention, video datasets monitored by different image acquisition devices at the second moment can be obtained, and the target cloud model can be trained using the acquired video acquisition devices. The trained target cloud model is updated through the cloud-edge model, and the model in the edge is updated to obtain the target edge model.
[0071] Optionally, the target edge model can be updated to a model obtained by updating the original edge model based on the target edge model, thereby obtaining the target edge model. The updated target edge model can be used to identify the video dataset monitored by different image acquisition devices at the third time.
[0072] In this embodiment, assuming the original edge model is the initial model, it can be trained on the edge using video datasets monitored by different image acquisition devices at the first moment. The original edge model trained in the customer's online environment is then used to update the model in the cloud, resulting in the original cloud model. The cloud then uses the original cloud model to identify the input images in the scene to be monitored, obtaining feature vectors corresponding to each scene. These feature vectors are processed and used as supervisory information to train the target cloud model. The target cloud model is then trained on video datasets monitored by different image acquisition devices at the second moment, resulting in the target edge model. The target edge model can be updated to the model obtained after updating the original edge model with the target edge model at the edge. This target edge model can then identify video datasets monitored by different image acquisition devices at the third moment. This embodiment of the invention achieves collaborative updating of the cloud-edge model through the above method, thereby improving the model update efficiency.
[0073] Through steps S202 to S208 of this application, an original cloud model from the edge is obtained. This original cloud model is obtained by training an original edge model using a video dataset, which is obtained by the edge using different image acquisition devices monitored at a first moment. Image recognition is performed on the input image in the scene to be monitored based on the original cloud model to obtain feature vectors. A target cloud model is then trained based on these feature vectors. Finally, the target cloud model is trained using video datasets monitored at a second moment by different image acquisition devices to obtain a target edge model. This target edge model is used to update the original edge model, enabling the edge to recognize video datasets monitored by different image acquisition devices at a third moment based on the target edge model. In other words, this embodiment of the invention, without manual annotation, obtains an original cloud model by training the original edge model using the video dataset monitored at a first moment, and obtains a target edge model by training the target cloud model using the video dataset monitored at a second moment. By coordinating the updates of the original cloud model and the original edge model using a large amount of video data, the technical effect of timely model updates is achieved, solving the technical problem of low model update efficiency.
[0074] The method described in this embodiment will be further described below.
[0075] As an optional implementation, step S204, based on the original cloud model, performs image recognition on the input image in the scene to be monitored to obtain feature vectors, including: based on the original cloud model corresponding to the multiple scenes to be monitored, performing image recognition on the input image in the multiple scenes to be monitored to obtain multiple sub-feature vectors, wherein the multiple sub-feature vectors correspond one-to-one with the multiple scenes to be monitored; and performing average pooling on the multiple sub-feature vectors to obtain feature vectors.
[0076] In this embodiment, the original cloud model can correspond to multiple monitoring scenarios. Based on the original cloud model corresponding to each of the multiple monitoring scenarios, the input images in the multiple monitoring scenarios are identified to obtain multiple sub-feature vectors that correspond one-to-one with the multiple monitoring scenarios. The multiple sub-feature vectors are then subjected to average pooling to obtain feature vectors. The monitoring scenario can be a scenario preset according to actual needs; the feature vectors can be used to represent the image.
[0077] Optionally, each scene to be monitored can correspond to a scene model. Input images from multiple scenes to be monitored are obtained, and the scene model is used to identify the input images to obtain multiple sub-feature vectors corresponding to the multiple scenes to be monitored. The multiple sub-feature vectors are then averaged and fused to obtain a feature vector.
[0078] For example, identification can be directly performed on unlabeled video datasets. Multiple scene models are used to process the images to be monitored in the video dataset, obtaining sub-feature vectors corresponding to the input image in each scene. These sub-feature vectors can be further processed using F1, F2, ..., F... n This means that by performing average pooling (i.e., fusion) on multiple sub-feature vectors, a feature vector F representing the image to be monitored is obtained. This feature vector can then be calculated using the following formula:
[0079]
[0080] As an optional implementation, step S206, training the target cloud model based on the feature vector, includes: training the original cloud model based on the feature vector and the target loss function to obtain the target cloud model.
[0081] In this embodiment, the feature vector can be used as supervision information to train the original cloud model based on the target loss function to obtain the target cloud model. The target loss function can be the relative entropy loss function (Kullback-Leibler divergence, or KL divergence for short) and the adversarial loss. The adversarial loss can be composed of a simple discriminant network, which can be used to help the data generated by the target cloud model and the feature vector belong to the same distribution, thereby improving the accuracy of the model prediction.
[0082] Optionally, the image to be recognized can be processed by multiple scene models (which can be reflow models) to obtain the sub-feature vectors corresponding to the input image in each scene. The sub-feature vectors in each scene are obtained, and the multiple sub-feature vectors are averaged (i.e. fused) to obtain the feature vector representing the input image. The feature vector F is used as supervision information. Using the target loss function, a discriminant network is constructed to help the network fit the distribution of other models as much as possible to adjust the parameters of the original cloud model and obtain the target cloud model.
[0083] For example, suppose there are four city scene reflow models (or model reflows) namely Scene 1 model (M1), Scene 2 model (M2), Scene 3 model (M3), and Scene 4 model (M4). The model to be trained is M0. All input images are processed by the four city scene models (four city scene reflow models), generating four sub-feature vectors (F1, F2, F3, F4). The four sub-feature vectors are average pooled (i.e. fused) to obtain the feature vector F representing the input image. The original cloud model is trained using the target loss function so that its output is a feature vector close to the original feature vector F, thus obtaining the target cloud model.
[0084] As an optional implementation, step S208 involves training the target cloud model using video datasets monitored by different image acquisition devices at the second moment to obtain the target edge model. This includes: identifying the monitored video dataset based on the original edge model to obtain identification results; determining the identification results as pseudo-labels and training the target cloud model based on the pseudo-labels; or training the target cloud model based on the pseudo-labels of the monitored video dataset to obtain the target edge model.
[0085] In this embodiment, the pseudo-labels are obtained by the edge-end through the original edge-end model identifying the monitored video dataset. The identification results are then used as pseudo-labels. The target cloud model can be trained based on these pseudo-labels to obtain the target cloud model, or the target cloud model can be trained based on the pseudo-labels of the monitored video dataset. The target cloud model is then updated in the edge-end through model updates to obtain the target edge model. The pseudo-labels can be pre-annotated labels obtained through inference from the original edge-end model, and can be represented by P... O The term "to perform" is used here for illustrative purposes only and is not intended to impose any specific limitations.
[0086] Optionally, the video datasets monitored by different image acquisition devices at the second moment can be obtained. The data in the video dataset can be sent to the quality evaluation module frame by frame in sequence, and images that do not meet the requirements can be filtered out to obtain the input image. The pseudo-label of the input image obtained by the edge recognition through the original edge model is determined as the pseudo-label of the video dataset. The original edge model may include a multi-classification model.
[0087] For example, it can be assumed that the original edge model includes a multi-classification model. The pseudo-labels mentioned above can be determined by the classification model in the edge after monitoring the input car image. The target cloud model is trained based on the pseudo-labels of the monitored car image, and the original edge model is updated based on the trained target cloud model to obtain the target edge model.
[0088] As an optional implementation method, the target cloud model is trained based on the pseudo-labels of the monitored video dataset to obtain the target edge model, including: supervised training of the target cloud model based on the pseudo-labels of the monitored video dataset to obtain the target edge model.
[0089] In this embodiment, pseudo-labels identified by the original edge model are obtained, supervised training of the target cloud model is performed based on the pseudo-labels, and the original edge model is updated based on the trained target cloud model to obtain the target edge model.
[0090] Alternatively, a method based on Contrastive Language-Image Pre-training (CLIP) can be used. This method uses an image encoder to recognize and generate a descriptive language segment from the video dataset. At the same time, a natural language encoder is used to process the pseudo-labels to obtain another descriptive language segment. The image encoder and the natural language encoder are then trained based on these two descriptive languages to obtain a small edge model+ that can represent the data distribution.
[0091] In this embodiment of the invention, a training framework based on the clip paradigm is proposed. By constructing "image-natural language pairs" through the "label-natural language" in the model, the image encoder and the natural language encoder are trained. This allows the vectors between the image encoder and the natural language encoder to support each other, narrowing the distance between them and completing the image-natural language comparative training. By using the clip paradigm-based training framework in the original cloud model, and using language as supervised information to train image features, the single-modal image data is expanded into multimodal data, thereby improving the accuracy of model prediction.
[0092] As an optional implementation, step S208 involves training the target cloud model using video datasets monitored by different image acquisition devices at the second moment to obtain the target edge model, including: performing transfer learning on the weights of the target cloud model based on the monitored video dataset to obtain the target edge model.
[0093] In this embodiment, a target cloud model is trained based on pseudo-labeled data, and the weights (network parameters) of the target cloud model are transferred to obtain the target edge model.
[0094] Optionally, the target cloud model can be iteratively optimized based on the identified pseudo-labels using the above method to train a powerful and excellent network weight. The weights of the target cloud model obtained after training can be transferred and adapted to obtain the target edge model.
[0095] For example, after obtaining a powerful network weight based on pseudo-label training, the target cloud model (large cloud model) is used to select data from the target target domain to perform transfer learning on the target edge model. That is, the target cloud model and the customer's online environment (data from the target domain) are used for transfer adaptation. The learned large cloud model replaces the small edge model on the online platform to obtain the target edge model, thereby completing the update of the model in the customer's online edge.
[0096] In the cloud-side embodiment of the present invention, without the need for manual annotation, the original edge model is trained using the video dataset monitored at the first moment to obtain the original cloud model, and the target cloud model is trained using the video dataset monitored at the second moment to obtain the target edge model. Thus, by using a large amount of video data to coordinate and update the original cloud model and the original edge model, the technical effect of timely model updating is achieved, solving the technical problem of low efficiency in model updating.
[0097] The image processing method of this invention will now be described from the perspective of the edge.
[0098] Figure 3 This is a flowchart of another image processing method according to an embodiment of the present invention, such as... Figure 3 As shown, the method may include the following steps:
[0099] Step S302: Obtain the video dataset monitored by different image acquisition devices at the first moment.
[0100] In the technical solution provided by step S302 of the present invention, different image acquisition devices are deployed in the scene to be monitored, and the edge acquires the video dataset monitored by the different image acquisition devices at the first moment. The video dataset can be streaming data.
[0101] For example, by building a streaming dataloader, edge-side small models can continuously obtain streaming data. Streaming data can include common message queues (RocketMQ), production-level consumer queues (Kafka), search engines (Elasticsearch), etc. Different data extraction interfaces can be designed for different sources. It should be noted that the types of streaming data mentioned here are only for illustrative purposes, and no specific restrictions are imposed on streaming data.
[0102] Step S304: The original edge model is trained using the monitored video dataset to obtain the original cloud model. The original cloud model is used to train the feature vector of the cloud based on the input image in the scene to be monitored, thereby obtaining the target cloud model.
[0103] In the technical solution provided in step S304 of the present invention, the edge can use the monitored video dataset to train the original edge model, and update the model in the cloud after training the original edge model to obtain the original cloud model. Based on the input image in the scene to be monitored, the feature vector is trained to obtain the target cloud model. The original cloud model can be a large cloud model in the running and production environment.
[0104] Optionally, the edge-side small model can be self-supervised pre-trained based on the acquired streaming data, and the trained original edge-side model can be used to update the model in the cloud to obtain the original cloud model.
[0105] Step S306: Obtain the video dataset monitored by different image acquisition devices in the cloud at the second moment to train the target cloud model, and obtain the target edge model.
[0106] In the technical solution provided by step S306 of the present invention, different image acquisition devices monitor the video dataset at the second moment, and use the acquired video dataset to train the original cloud model in the cloud to obtain the target edge model.
[0107] Step S308: Update the original edge model to the target edge model, wherein the target edge model is used to identify the video datasets monitored by different image acquisition devices at the third time.
[0108] In the technical solution provided by step S308 of the present invention, the original edge model trained in the customer's online environment is used to update the model in the cloud to obtain the original cloud model. The original cloud model is used to train in the cloud to obtain the target cloud model. The cloud model is used to train using video datasets monitored by different image acquisition devices at the second time to obtain the target cloud model. The trained target cloud model is used to update the model in the edge to obtain the target edge model. The updated target edge model can recognize video datasets monitored by different image acquisition devices at the third time.
[0109] Through steps S302 to S308 of the present invention, video datasets monitored by different image acquisition devices at the first moment are obtained; the original edge model is trained using the monitored video datasets to obtain the original cloud model, wherein the original cloud model is used to enable the cloud to train the feature vectors of the image based on the input image in the scene to be monitored, thereby obtaining the target cloud model; the target edge model is obtained by training the target cloud model using video datasets monitored by different image acquisition devices at the second moment; the original edge model is updated to the target edge model to identify the video datasets monitored by different image acquisition devices at the third moment, thereby achieving the technical effect of timely model updates and solving the technical problem of low efficiency in model updates.
[0110] The method described in this embodiment will be further described below.
[0111] As an optional implementation, step S308 updates the original edge model to the target edge model to identify the video datasets monitored by different image acquisition devices at the third time, including: determining the target edge model as the original edge model, determining the video dataset monitored at the third time as the video dataset monitored at the first time, and returning to execute the step of training the original edge model with the monitored video dataset to obtain the original cloud model.
[0112] In this embodiment, when the model in the edge needs to be updated again after a period of time, the model can be updated by following these steps: at this time, the target edge model that was previously trained can be determined as the original edge model, and the video dataset monitored at the third time moment can be determined as the video dataset monitored at the first time moment. The original edge model is trained using the monitored video dataset. By repeating the above steps of determining the target edge model, the goal of timely updating the target edge model within a predetermined time can be achieved.
[0113] For example, after a period of time, the target edge model will be determined as the original edge model, and the video dataset monitored at the third time will be determined as the video dataset monitored at the first time. Based on the monitored dataset, the original edge model will be trained again to obtain the target cloud model. The target cloud model will be trained using the video datasets monitored at the second time by different image acquisition devices to obtain the target edge model. The original edge model will be updated with the obtained target edge model to become the target edge model. The updated original edge model can then be used to identify the video datasets monitored at the third time by different image acquisition devices.
[0114] Optionally, the edge model can be updated again at target time intervals or when the video dataset reaches a certain amount of data using the above method. It should be noted that there are no restrictions on the triggering conditions for model updates here.
[0115] As an optional implementation, step S304, training the original edge model using the monitored video dataset to obtain the original cloud model, includes: performing self-supervised training on the original edge model using the monitored video dataset to obtain the original cloud model.
[0116] In this embodiment, the original edge model can be self-supervised and trained using the monitored video dataset to obtain the original cloud model.
[0117] Optionally, a self-supervised training module can be used to complete the self-supervised training of the original edge model, which may include a single-vision image paradigm and a multimodal paradigm.
[0118] As an optional implementation, the original edge model is self-supervised and trained using the monitored video dataset to obtain the original cloud model, including at least one of the following: augmenting the monitored video dataset to obtain an augmented video dataset; using the augmented video dataset as training samples to self-supervise and train the original edge model to obtain the original cloud model, wherein the distance between multiple positive samples in the training samples is less than the distance between multiple negative samples.
[0119] In this embodiment, the monitored video dataset can be augmented to obtain an augmented video dataset. The augmented video dataset can be used as training samples to perform self-supervised training on the original edge model, which can make the distance between multiple positive samples in the training samples smaller than the distance between multiple negative samples, thereby completing the self-supervised training of the original edge model. The trained original edge model is then used to update the model in the cloud to obtain the original cloud model.
[0120] Optionally, after performing different enhancements on the video dataset, the resulting enhanced video dataset is fed into the training network to reduce the distance between positive samples and increase the distance with other negative samples. Given image data x in the video dataset, the goal of contrastive learning is to learn an encoder f such that:
[0121] d(f(x), f(x) + ))<<d(f(x),f(x) - ))
[0122] Where, x + d is a positive sample similar to x; x- is a negative sample dissimilar to x; d() is a metric function used to measure the similarity between positive and negative samples, which can be Euclidean distance, cosine similarity, etc.
[0123] In this embodiment of the invention, the original edge model is trained by comparative learning in the edge to obtain the original cloud model.
[0124] As an optional implementation, the original edge model is self-supervised and trained using the monitored video dataset to obtain the original cloud model. This includes: masking the monitored video dataset to obtain a masked video dataset; restoring the masked video dataset based on the original edge model; and training the original edge model based on the restored masked video dataset and the monitored video dataset to obtain the original cloud model.
[0125] In this embodiment, the monitored video dataset can be masked to obtain a masked video dataset. The original edge model is then guided to restore the masked video dataset to obtain a restored masked video dataset. The original edge model is compared with the restored masked video dataset and the monitored video dataset to determine the loss function. Based on the loss function, the model parameters are adjusted to obtain the original cloud model.
[0126] Optionally, you can input the data to be processed (I input By masking certain areas, the original edge model is guided to restore the masked areas, resulting in output data (I). output Based on the output data and the data to be processed, a loss function is determined. The parameters of the original edge model are adjusted according to the loss function to achieve the purpose of training the original edge model. The loss function can be:
[0127] L = ||I input -I output ||
[0128] As an optional implementation, the original edge model is self-supervised and trained using the monitored video dataset to obtain the original cloud model, including: identifying pseudo-labels of the monitored video dataset based on the original edge model; generating text from the pseudo-labels of the monitored video dataset, wherein the text is used to describe the monitored video dataset; and training the original edge model into the original cloud model based on the monitored video dataset and the text.
[0129] In this embodiment, a video dataset can be input, and the original edge model monitors the video dataset to obtain multi-dimensional labels for the video dataset. These multi-dimensional labels are used to generate text in the cloud using a natural language encoder. At the same time, an image encoder can be used to process the video dataset to obtain the corresponding text for the video dataset. Based on the two obtained texts, the original edge model is trained to obtain the original cloud model. Here, the text can be natural descriptive language, and no specific restrictions are placed on the conversion of text information.
[0130] For example, an image of a car can be input into the original edge model, which can output the label "Vehicle color: black, vehicle brand: *g, vehicle view: front". The fields of the output label are processed by the image encoder and the natural language encoder to generate a natural language description: the front view of a black car, the vehicle brand is *g. The original cloud model is obtained through image-natural language comparison training.
[0131] In this embodiment of the invention, without the need for manual annotation, the original edge model is trained using the video dataset monitored at the first moment to obtain the original cloud model, and the target cloud model is trained using the video dataset monitored at the second moment to obtain the target edge model. Thus, by using a large amount of video data to coordinate and update the original cloud model and the original edge model, the technical effect of timely model updates is achieved, solving the technical problem of low efficiency in model updates.
[0132] As an optional implementation, the original edge model is trained into the original cloud model based on the monitored video dataset and the target text, including: comparing and training the original edge model based on the encoding vector of the monitored video dataset and the encoding vector of the target text to obtain the original cloud model.
[0133] In this embodiment, the CLIP paradigm can be used at the edge to generate encoding vectors based on the target text in the pseudo-labels of the monitored video dataset, and the video dataset can be processed to obtain the encoding vectors corresponding to the video dataset. The two encoding vectors are compared and trained. Through data updates between the customer's online scene and the generated scene, the trained model is used as the original cloud model. The encoding vectors can be used to describe objects in the image, and can be natural description language, symbols, etc. No specific restrictions are placed on the form of the encoding vectors here.
[0134] Optionally, a natural language encoder can be used to generate a natural language description from the pseudo-label fields. For example, the natural language description could be the front view of a black sedan, with the vehicle brand being *G. The natural language generation method can be set manually using a rule engine. For instance, for a vehicle classification task, assuming M... a The pseudo-labels output by the model are vehicle color - red, vehicle brand - *horse, vehicle model - x3, so the output natural language description is a red *horse x3.
[0135] Optionally, the original edge model can be trained using a self-supervised training method based on the language image pre-training paradigm. The model training can include image-natural language comparison training, which can include training the image encoder and the natural language encoder.
[0136] This invention also provides another image processing method that can be applied to traffic and road scenarios, such as urban monitoring scenarios, and the model can be used to identify videos monitored in urban traffic and road scenarios.
[0137] Figure 4 This is a flowchart of an image processing method according to an embodiment of the present invention, such as... Figure 4 As shown, the method may include the following steps.
[0138] Step S402: Obtain the original cloud model from the edge, wherein the original cloud model is obtained by training the original edge model with the video dataset, the video dataset is obtained by the edge using different image acquisition devices to monitor traffic roads at the first moment, and the video dataset contains at least one vehicle driving through the traffic road.
[0139] Step S404: Based on the original cloud model, identify the input image in the traffic road to obtain the feature vector;
[0140] Step S406: The target cloud model is obtained by training based on the feature vector.
[0141] Step S408: The target cloud model is trained using video datasets monitored by different image acquisition devices at the second time to obtain the target edge model. The target edge model is used to update the original edge model, so that the edge can identify the video datasets monitored by different image acquisition devices for traffic roads at the third time based on the target edge model.
[0142] Through steps S402 to S408 of the present invention, an original cloud model from the edge is obtained. The original cloud model is obtained by training an original edge model with a video dataset. The video dataset is obtained by the edge monitoring traffic roads at the first moment using different image acquisition devices. The video dataset contains at least one vehicle traveling through the traffic road. Based on the original cloud model, the input images in the traffic road are identified to obtain feature vectors. Based on the feature vectors, a target cloud model is trained. The target cloud model is trained with video datasets monitored by different image acquisition devices at the second moment to obtain a target edge model. The target edge model is used to update the original edge model, so that the edge can identify video datasets monitored by different image acquisition devices on the traffic road at the third moment based on the target edge model. Thus, by using a large amount of video data to coordinate the update of the original cloud model and the original edge model, the technical effect of timely model update is achieved, solving the technical problem of low efficiency in model update.
[0143] As another alternative embodiment, Figure 5 This is a schematic diagram of the hardware environment of a virtual reality device according to an embodiment of the present invention, which describes an image processing method. Figure 5 As shown, the virtual reality device 504 is connected to the terminal 506, and the terminal 506 is connected to the server 502 via a network. The virtual reality device 504 is not limited to: virtual reality headsets, virtual reality glasses, virtual reality all-in-one machines, etc. The terminal 506 is not limited to PCs, mobile phones, tablets, etc. The server 502 can be a server corresponding to a media file operator. The network mentioned above includes, but is not limited to: wide area network, metropolitan area network, or local area network.
[0144] Optionally, the virtual reality device 504 in this embodiment includes a memory, a processor, and a transmission device. The memory stores an application program that can be used to execute: acquiring an original cloud model from the edge, wherein the original cloud model is obtained by training an original edge model with a video dataset, the video dataset being obtained by the edge using different image acquisition devices at a first moment; recognizing input images in the scene to be monitored based on the original cloud model to obtain feature vectors; training a target cloud model based on the feature vectors; training the target cloud model with video datasets monitored by different image acquisition devices at a second moment to obtain a target edge model, wherein the target edge model is used to update the original edge model, enabling the edge to recognize video datasets monitored by different image acquisition devices at a third moment based on the target edge model, achieving the technical effect of timely model updates and solving the technical problem of low efficiency in model updates.
[0145] The terminal in this embodiment can be used to: display the monitored video dataset on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device; display the input image in the scene to be monitored on the VR or AR device; identify the input image based on the original cloud model from the edge to obtain feature vectors; train a target cloud model based on the feature vectors, wherein the original cloud model is obtained by training the original edge model with the video dataset, and the video dataset is obtained by the edge using different image acquisition devices at the first moment; train the target cloud model with the video datasets monitored by different image acquisition devices at the second moment to obtain a target edge model, wherein the target edge model is used to update the original edge model, so that the edge can identify the video datasets monitored by different image acquisition devices at the third moment based on the target edge model; drive the VR or AR device to display the video dataset monitored at the third moment, and the virtual reality device 504 displays it at the target projection position after receiving the monitored video dataset.
[0146] Optionally, the eye-tracking head-mounted display (HMD) and eye-tracking module in the virtual reality device 504 of this embodiment function the same as in the embodiments described above. That is, the screen in the HMD is used to display real-time images, and the eye-tracking module in the HMD is used to acquire the real-time movement path of the user's eyes. The terminal in this embodiment acquires the user's position and movement information in real three-dimensional space through the tracking system, and calculates the three-dimensional coordinates of the user's head in virtual three-dimensional space, as well as the user's field of vision orientation in virtual three-dimensional space.
[0147] Figure 5 The hardware structure block diagram shown can serve not only as an exemplary block diagram of the aforementioned AR / VR device (or mobile device), but also as an exemplary block diagram of the aforementioned server. In the operating environment described above, the present invention also provides, for example... Figure 6 The image processing method shown can be applied to virtual reality (VR) devices or augmented reality (AR) devices, and the model can be used to analyze video segments in VR or AR devices. It should be noted that the image processing method in this embodiment can be... Figure 6 The mobile terminal of the illustrated embodiment is executed.
[0148] Figure 6 This is a flowchart of another image processing method according to an embodiment of the present invention, such as... Figure 6 As shown, the method may include the following steps.
[0149] Step S602: Display the input image of the scene to be monitored on the presentation screen of the virtual reality (VR) device or augmented reality (AR) device.
[0150] In step S604, the virtual display (VR) device or augmented reality (AR) device sends the input image to the edge. The original cloud model of the edge model performs image recognition on the input image and trains the target cloud model based on the recognized feature vector. The original cloud model is obtained by training the original edge model with a video dataset. The video dataset is obtained by the edge using different image acquisition devices to monitor the image at the first moment.
[0151] Step S608: After training the target cloud model with video datasets monitored by different image acquisition devices at the second moment, and updating the original edge model based on the trained target edge model, the VR device or AR device is driven to render and display the video dataset monitored at the third moment.
[0152] Optionally, in this embodiment, the image processing method described above can be applied to a hardware environment consisting of a server and a virtual reality device. The video is displayed on the screen of the virtual reality device or augmented reality device. The server can be a server corresponding to a media file operator. The aforementioned network includes, but is not limited to, a wide area network (WAN), a metropolitan area network (MAN), or a local area network (LAN). The aforementioned virtual reality device is not limited to, for example, a virtual reality headset, virtual reality glasses, or a standalone virtual reality device.
[0153] It should be noted that the image processing method described above in this embodiment, when applied to VR or AR devices, may include... Figure 6 The method of the illustrated embodiment is used to drive a VR device or AR device to display a video dataset monitored in a third moment.
[0154] Optionally, the processor in this embodiment can invoke the application stored in the memory via the transmission device to perform the above steps. The transmission device can receive media files sent by the server via a network, and can also be used for data transmission between the processor and the memory.
[0155] Optionally, in a virtual reality device, there is a head-mounted display with eye tracking. The screen in the HMD is used to display the video footage. The eye-tracking module in the HMD is used to acquire the real-time movement path of the user's eyes. The tracking system is used to track the user's position and movement information in real three-dimensional space. The computing and processing unit is used to acquire the user's real-time position and movement information from the tracking system and calculate the three-dimensional coordinates of the user's head in the virtual three-dimensional space, as well as the user's field of vision orientation in the virtual three-dimensional space.
[0156] In this embodiment of the invention, the virtual reality device can be connected to a terminal, and the terminal and the server are connected through a network. The virtual reality device is not limited to: virtual reality helmet, virtual reality glasses, virtual reality all-in-one machine, etc. The terminal is not limited to PC, mobile phone, tablet computer, etc. The server can be the server corresponding to the media file operator. The network includes, but is not limited to: wide area network, metropolitan area network or local area network.
[0157] Figure 7 This is a schematic diagram of another image processing result according to an embodiment of the present invention, such as... Figure 7 As shown, this drives VR or AR devices to display video datasets monitored in the third moment.
[0158] This invention displays an input image of a scene to be monitored on the display screen of a virtual reality (VR) device or an augmented reality (AR) device. The VR or AR device sends the input image to the edge, where the original cloud model performs image recognition on the input image and trains a target cloud model based on the recognized feature vectors. The original cloud model is obtained by training the original edge model with a video dataset, which is obtained by the edge using different image acquisition devices to monitor the scene at a first moment. After training the target cloud model with the video dataset monitored by different image acquisition devices at a second moment and updating the original edge model based on the trained target edge model, the VR or AR device is driven to render and display the video dataset monitored at a third moment. By utilizing a large amount of video data to coordinate the update of the original cloud model and the original edge model, the technical effect of timely model updates is achieved, solving the technical problem of low efficiency in model updates.
[0159] Example 2
[0160] The preferred implementation of the method described above in this embodiment will be further described below, specifically using a method of co-evolution of large and small models based on self-supervision and multi-model distillation.
[0161] In urban scenarios, tens of thousands of video points can generate a large amount of video data every day. How to fully perceive and understand this data is a very important step in recognition algorithms for urban scenarios.
[0162] In related technologies, urban scene monitoring data is often stored on the client's intranet, which has great privacy. Directly labeling video data before training has the problem of excessively long cycle, as well as certain data privacy ethics risks and leakage risks. At the same time, due to the large amount of data acquired, the model often cannot be updated in a timely manner, which also has a certain lag.
[0163] To address the aforementioned issues and enable the full utilization of massive urban scene data for model updates, this invention proposes a large-scale model co-evolution system based on self-supervised and multi-model distillation. This system includes an edge-based self-supervised training module and a cloud-based multi-model distillation module. The cloud-based multi-model distillation module further includes an edge-based iterative optimization module. These three modules enable the full utilization of massive data to co-evolve the large cloud-based model and the small edge-based model for visual tasks without manual annotation, thereby achieving timely model updates and resolving the problem of low model update efficiency.
[0164] The following section provides a further introduction to the edge model self-supervised training module applied to the customer's online environment.
[0165] The edge model self-supervised module constructs a streaming data loader, enabling the edge mini-model to continuously acquire streaming data. Based on the acquired streaming data, the edge mini-model is self-supervised pre-trained to obtain the final edge mini-model+. The streaming data can include common message middleware (RocketMQ), production-level consumer middleware (Kafka), search engines (Elasticsearch), etc. Different data extraction interfaces can be designed for different sources. It should be noted that the types of streaming data mentioned here are only for illustrative purposes, and no specific restrictions are imposed on the streaming data.
[0166] As an alternative implementation, training paradigms for edge-based small models can include single-vision image paradigms and multimodal paradigms.
[0167] In this embodiment, the single-vision paradigm may include contrastive learning and generative learning.
[0168] Optionally, Figure 8 This is a schematic diagram of a contrastive learning method according to an embodiment of the present invention, such as... Figure 8 As shown, contrastive learning involves comparing data with both positive and negative samples in the training network to learn the feature representations of the samples.
[0169] Optionally, the input image can be augmented in different ways before being fed into the training network to reduce the distance between positive samples and increase the distance with other negative samples. Given image data x, the goal of contrastive learning is to learn an encoder f such that:
[0170] d(f(x), f(x) + ))<<d(f(x),f(x) - ))
[0171] Where, x + d is a positive sample similar to x; x- is a negative sample dissimilar to x; d() is a metric function used to measure the similarity between positive and negative samples, which can be Euclidean distance, cosine similarity, etc.
[0172] Optionally, generative learning can take as input data to be processed (I input By masking certain areas, the network is guided to restore the masked areas, resulting in output data (Io). utput This achieves the goal of training the network, and its loss function is:
[0173] L = ||I input -I output ||
[0174] In this embodiment, the multimodal paradigm can employ a method based on contrastive language-image pre-training (CLIP), using language as supervisory information to train the features of the images. Figure 9 This is a schematic diagram of a self-supervised training framework based on a language image pre-training paradigm according to an embodiment of the present invention, as shown below. Figure 9 As shown, by inputting a small image of a car, and processing it based on the language image pre-training paradigm, labels in multiple dimensions can be obtained.
[0175] For example, such as Figure 9 As shown, we can assume that we already have model M. a (This can be a multi-class classification model). Input a car image and output the label "Vehicle color: black, vehicle brand: *g, vehicle view: front". The image encoder and natural language encoder in the CLIP paradigm process the fields of the output label to generate a natural description: the front view of a black sedan, the vehicle brand is *g.
[0176] The following section further introduces the self-supervised training method for edge models based on the language-image pre-training paradigm in the multimodal paradigm.
[0177] Figure 10 This is a flowchart of a self-supervised training method for edge models based on a language image pre-training paradigm according to an embodiment of the present invention, as follows: Figure 10 As shown, the edge model self-supervised training method based on the language image pre-training paradigm can include the following steps:
[0178] Step S1002: Obtain the image of the car.
[0179] In this embodiment, streaming data (dataloader) is continuously acquired to obtain car images. These car images can be sent sequentially to the quality evaluation module to filter out a large number of small images that do not meet the requirements and reduce a large number of small images with no information.
[0180] Step S1004: Process the car image using a multi-classification model.
[0181] In this embodiment, the selected car image (Img) x The data is fed into a multi-classification model, which outputs multiple labels. The labels output by the model are used as pseudo-labels for the car image.
[0182] Step S1006: Process the output results of the multi-classification model.
[0183] In this embodiment, a natural description language can be generated based on the fields of the pseudo-tag using the CLIP paradigm. For example, the natural description language could be the front view of a black sedan with the vehicle brand being *G.
[0184] Alternatively, the above-mentioned natural language generation method can be set using a manually configured rule engine. For example, for a vehicle classification task, suppose M... a The pseudo-labels output by the model are vehicle color - red, vehicle brand - *horse, vehicle model - x3, so the output natural language description is a red *horse x3.
[0185] In this embodiment, the model can be trained using a self-supervised training method based on the language image pre-training paradigm. The model training can include image-natural language comparison training, which can include training the image encoder and the natural language encoder.
[0186] Optionally, the vectors between the image encoder and the natural language encoder can be made to support each other, thus narrowing the distance between them and completing the image-natural language contrast training.
[0187] In this embodiment of the invention, a small edge model+ that can characterize the data distribution can be obtained through the self-supervised training process of the edge model.
[0188] The following section provides a further introduction to the cloud-based multi-model distillation module, which can be applied to the company's production environment.
[0189] In this embodiment, the multi-model distillation module can be used to better fit the data distribution in multiple scenarios (e.g., urban scenarios such as scenario 1, scenario 2, scenario 3... scenario N), thereby improving the generalization ability of the model.
[0190] Figure 11 This is a schematic diagram of a cloud-based multi-model distillation module according to an embodiment of the present invention, as shown below. Figure 11 As shown, the input image is processed under different scene models to obtain feature vectors. All feature vectors of the four scenes are averaged and pooled to obtain feature vectors representing the image. The fused feature vectors are used as supervision information to conduct supervised training of the network to obtain a cloud-based large model+ (which can be the target cloud model).
[0191] Figure 12 This is a flowchart of a cloud-based multi-model distillation method according to an embodiment of the present invention, such as... Figure 12 As shown, the cloud-based multi-model distillation method may include the following steps:
[0192] Step S1202: Perform multi-model inference processing on the input image.
[0193] In this embodiment, unlabeled data can be directly used to train the cloud-based multi-model distillation method. Multiple scene models are used to process the image to be recognized to obtain the feature vector corresponding to the input image in each scene.
[0194] Step S1204: Obtain the feature vector for each scene and fuse multiple feature vectors.
[0195] In this embodiment, feature vectors for each scene are obtained, and average pooling (i.e., fusion) is performed on multiple feature vectors to obtain a vector F representing the sample, which can be calculated using the following formula:
[0196]
[0197] For example, such as Figure 12As shown, assume there are four city scene reflow models (or model reflows) namely Scene 1 model (M1), Scene 2 model (M2), Scene 3 model (M3), and Scene 4 model (M4). The model to be trained is M0. All input images are processed by the four city scene models (four city scene reflow models) to generate four feature vectors (F1, F2, F3, F4). The four feature vectors are then averaged (i.e. fused) to obtain the vector F representing the sample.
[0198] Step S1206: Use the feature vector F as supervision information for supervised training.
[0199] In this embodiment, the feature vector F is used as supervision information to perform supervised training on the multi-model distillation framework. The loss function in the supervision process can be the relative entropy loss function plus the adversarial loss. The adversarial loss can be composed of a simple discriminant network to help the data generated by the network to belong to the same distribution as the original vector F.
[0200] Optionally, network M (multi-model distillation framework) is trained to output a feature vector that is close to the original feature vector F. The loss function can be the KL loss function. At the same time, a discriminant network is constructed to help the network fit the distribution of other models as much as possible, thereby obtaining the cloud-based large model+.
[0201] The following section provides a further introduction to the edge iteration optimization module, which can be applied to the customer's online environment.
[0202] Figure 13 This is a schematic diagram of an edge model iterative optimization method according to an embodiment of the present invention, as shown below. Figure 13 As shown, edge model iteration can include the following two methods.
[0203] As an alternative embodiment, such as Figure 13 As shown, after obtaining a powerful network weight based on pseudo-labels, the cloud-based large model can be used to select streaming data in the target domain for transfer learning. This completes the transfer adaptation based on the existing model and the target domain data. The learned cloud-based large model is then used to replace the online edge-based small model, resulting in the final edge-based small model, thus completing the update of the edge-based small model.
[0204] As an alternative implementation, the edge-based small model can be fine-tuned and trained on streaming data to obtain the large model in the cloud.
[0205] Alternatively, assuming the original edge-side small model is M0, a large cloud model (Mn) is obtained through fine-tuning and training, where M... NIt can include a backbone network and a head network, and the weights of the backbone part can be obtained through multi-model distillation.
[0206] Optionally, for streaming data in online tasks, pseudo-labels P can be obtained through the original edge-side small model. O Using pseudo-labels P O Supervised training is performed on the large cloud model until it passes the automatic evaluation criteria of the customer's online environment, resulting in the large cloud model+. The trained large cloud model is then used to replace the small edge model, thereby completing the update of the cloud-edge model.
[0207] It should be noted that the iterative optimization of the edge small model+ can be achieved by directly using the weights of the cloud large model as pre-trained weights and then performing transfer learning offline according to a specific task, or by using the weights of the cloud large model as fixed weights and using the pseudo-labels generated by the original edge small model for training. Either method can be chosen.
[0208] In this embodiment of the invention, a co-evolutionary system for large and small models based on self-supervised (edge model self-supervised training) and multi-model distillation (cloud multi-model distillation) is provided. The system includes three modules: an edge model self-supervised training module, a cloud multi-model distillation module, and an edge iterative optimization module. In the edge model self-supervised training module, a training framework based on the clip paradigm is proposed. By constructing image-natural language pairs through model-label-natural language, single-modal image data is expanded into multimodal data. The cloud multi-model distillation module can perform distillation operations on large cloud models obtained from multiple city scenes to extract a more robust large cloud model. This embodiment of the invention achieves the co-evolution of large cloud models and small edge models for visual tasks by making full use of massive data without manual annotation through the above three modules.
[0209] In another alternative embodiment, Figure 14 A block diagram illustrates one embodiment of using the computer terminal 10 (or mobile device) shown in Figure 1 above as a computing node in computing environment 1401. Figure 14 This is a structural block diagram of a computing environment according to an embodiment of the present invention, such as... Figure 14As shown, computing environment 1401 includes multiple compute nodes (such as servers) running on a distributed network (represented as 1410-1, 1410-2, ..., in the diagram). Each compute node contains local processing and memory resources, and end user 1402 can remotely run applications or store data within computing environment 1401. Applications can be provided as multiple services 1420-1, 1420-2, 1420-3, and 1420-4 within computing environment 1401, representing services "A", "D", "E", and "H", respectively.
[0210] End user 1402 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or requests of end user 1402 can be provided to ingress gateway 1430. Ingress gateway 1430 may include a corresponding agent to handle provisioning and / or requests for service 1420 (one or more services provided in computing environment 1401).
[0211] Service 1420 is provided or deployed based on various virtualization technologies supported by computing environment 1401. In some embodiments, service 1420 may be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. Virtual machine-based virtualization may involve simulating a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While the machine is virtualized by a virtual machine, container-based virtualization may launch containers to virtualize an entire operating system (OS), allowing multiple workloads to run on a single OS instance.
[0212] In one embodiment based on container virtualization, several containers of service 1420 can be assembled into a POD (e.g., a Kubernetes POD). For example, such as Figure 14 As shown, service 1420-2 can be equipped with one or more PODs 1440-1, 1440-2, ..., 1440-N (collectively referred to as POD 1440). Each POD 1440 can include a proxy 1445 and one or more containers 1442-1, 1442-2, ..., 1442-M (collectively referred to as container 1442). One or more containers 1442 in POD 1440 handle requests related to one or more corresponding functions of the service, and the proxy 1445 typically controls service-related network functions such as routing and load balancing. Other services 1420 can also be accompanied by PODs similar to POD 1440.
[0213] During operation, executing a user request from end user 1402 may require calling one or more services 1420 in computing environment 1401. Executing one or more functions of one service 1420 requires calling one or more functions of another service 1420. For example... Figure 14 As shown, service "A" 1420-1 receives user requests from terminal user 1402 from ingress gateway 1430. Service "A" 1420-1 can call service "D" 1420-2, and service "D" 1420-2 can request service "E" 1420-3 to perform one or more functions.
[0214] The aforementioned computing environment can be a cloud computing environment, where resource allocation is managed by cloud services, allowing functionality development without needing to consider implementation, adjustment, or server scaling. This computing environment allows developers to execute event-responsive code without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically and independently scale, rather than scaling a single hardware device to handle potential loads.
[0215] In another alternative embodiment, Figure 15 A block diagram illustrates one embodiment of using the computer terminal 10 (or mobile device) shown in Figure 1 above as a service mesh. Figure 15 This is a structural block diagram of a service mesh for an image processing method according to an embodiment of the present invention, such as... Figure 15 As shown, the Service Mesh 1500 is mainly used to facilitate secure and reliable communication between multiple microservices. Microservices refer to the decomposition of an application into multiple smaller services or instances, which are distributed across different clusters / machines.
[0216] like Figure 15 As shown, a microservice may include application service instance A and application service instance B, which together form the functional application layer of service mesh 1500. In one implementation, application service instance A runs as a container / process 1508 on machine / workload container group 1514 (POD), and application service instance B runs as a container / process 1510 on machine / workload container group 1516 (POD).
[0217] In one implementation, application service instance A can be a product query service, and application service instance B can be a product order placement service.
[0218] like Figure 15As shown, application service instance A and grid agent (sidecar) 1503 coexist in machine workload container group 614, and application service instance B and grid agent 1505 coexist in machine workload container 1514. Grid agents 1503 and 1505 form the data plane layer of service mesh 1500. Grid agents 1503 and 1505 run as container / process 1504, which can receive requests 1512 for product query services, and as grid agent 1506. Grid agent 1503 and application service instance A can communicate bidirectionally, and grid agent 1505 and application service instance B can also communicate bidirectionally. Furthermore, grid agents 1503 and 1505 can also communicate bidirectionally with each other.
[0219] In one implementation, all traffic from application service instance A is routed to the appropriate destination via mesh proxy 1503, and all network traffic from application service instance B is routed to the appropriate destination via mesh proxy 1505. It should be noted that the network traffic mentioned here includes, but is not limited to, Hypertext Transfer Protocol (HTTP), Representational State Transfer (REST), high-performance, general-purpose open-source frameworks (gRPC), and open-source in-memory data structure storage systems (Redis).
[0220] In one implementation, the functionality of the extended data plane layer can be achieved by writing custom filters for the agents (Envoy) in service mesh 1500. The service mesh agent configuration can enable the service mesh to correctly proxy service traffic, achieving service interoperability and service governance. Mesh agents 1503 and 1505 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.
[0221] like Figure 15 As shown, the service mesh 1500 also includes a control plane layer. This control plane layer can consist of a set of services running in a dedicated namespace, hosted by a managed control plane component 1501 within machine / workload container groups (machine / Pods) 1502. For example... Figure 15 As shown, the managed control plane component 1501 communicates bidirectionally with grid agents 1503 and 1505. The managed control plane component 1501 is configured to perform several control and management functions. For example, the managed control plane component 1501 receives telemetry data transmitted by grid agents 1503 and 1505 and can further aggregate this telemetry data. In addition to these services, the managed control plane component 1501 can also provide a user-facing application programming interface (API) to facilitate manipulation of network behavior and provision of configuration data to grid agents 1503 and 1505. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0222] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0223] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0224] Example 3
[0225] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 3 The image processing apparatus shown is an image processing method.
[0226] Figure 16 This is a schematic diagram of an image processing apparatus according to an embodiment of the present invention. Figure 16As shown, the image processing device 1600 may include: a first acquisition unit 1602, a first recognition unit 1604, a first training unit 1606, and a second training unit 1608.
[0227] The first acquisition unit 1602 is used to acquire the original cloud model from the edge, wherein the original cloud model is obtained by training the original edge model with the video dataset, and the video dataset is obtained by the edge using different image acquisition devices to monitor at the first moment.
[0228] The first recognition unit 1604 is used to perform image recognition on the input image in the scene to be monitored based on the original cloud model, and obtain the feature vector.
[0229] The first training unit 1606 is used to train the target cloud model based on feature vectors.
[0230] The second training unit 1608 is used to train the target cloud model using video datasets monitored by different image acquisition devices at the second time, so as to obtain the target edge model. The target edge model is used to update the original edge model, so that the edge can identify the video datasets monitored by different image acquisition devices at the third time based on the target edge model.
[0231] It should be noted that the first acquisition unit 1602, the first identification unit 1604, the first training unit 1606, and the second training unit 1608 mentioned above correspond to steps S302 to S308 in Embodiment 1. The four units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above units, as part of the device, can run in the computer terminal A provided in Embodiment 1.
[0232] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 4 The image processing apparatus shown is an image processing method.
[0233] Figure 17 This is a schematic diagram of an image processing apparatus according to an embodiment of the present invention, which is applied to an edge method side. Figure 17 As shown, the image processing device 1700 may include: a second acquisition unit 1702, a third training unit 1704, a third acquisition unit 1706, and an update unit 1708.
[0234] The second acquisition unit 1702 is used to acquire the video dataset monitored by different image acquisition devices at the first moment.
[0235] The third training unit 1704 is used to train the original edge model using the monitored video dataset to obtain the original cloud model. The original cloud model is used to train the feature vector of the cloud based on the input image in the scene to be monitored, thereby obtaining the target cloud model.
[0236] The third acquisition unit 1706 is used to acquire video datasets monitored by different image acquisition devices in the cloud at the second moment to train the target cloud model and obtain the target edge model.
[0237] Update unit 1708 is used to update the original edge model to the target edge model, wherein the target edge model is used to identify the video datasets monitored by different image acquisition devices at the third time.
[0238] It should be noted that the second acquisition unit 1702, the third training unit 1704, the third acquisition unit 1706, and the update unit 1708 mentioned above correspond to steps S402 to S408 in Embodiment 1. The four units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units, as part of the device, can run in the computer terminal A provided in Embodiment 1.
[0239] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 5 The image processing apparatus shown is applicable to traffic scenarios, such as urban monitoring. The model can be used to identify vehicles detected in urban traffic videos. No specific limitations are imposed on the objects that can be identified.
[0240] Figure 18 This is a schematic diagram of another image processing apparatus according to an embodiment of the present invention. Figure 18 As shown, the image processing device 1800 may include: a third acquisition unit 1802, a second recognition unit 1804, a third training unit 1806, and a fourth training unit 1808.
[0241] The third acquisition unit 1802 is used to acquire the original cloud model from the edge, wherein the original cloud model is obtained by training the original edge model with video dataset, the video dataset is obtained by the edge using different image acquisition devices to monitor traffic roads at the first moment, and the video dataset contains at least one vehicle driving through the traffic road.
[0242] The second recognition unit 1804 is used to recognize the input image in the traffic road based on the original cloud model and obtain the feature vector.
[0243] The third training unit 1806 is used to train the target cloud model based on feature vectors.
[0244] The fourth training unit 1808 is used to train the target cloud model using video datasets monitored by different image acquisition devices at the second time, to obtain the target edge model. The target edge model is used to update the original edge model, so that the edge can recognize the video datasets of traffic roads monitored by different image acquisition devices at the third time based on the target edge model.
[0245] It should be noted that the third acquisition unit 1802, the second identification unit 1804, the third training unit 1806, and the fourth training unit 1808 mentioned above correspond to steps S502 to S508 in Embodiment 1. The four units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above units, as part of the device, can run in the computer terminal A provided in Embodiment 1.
[0246] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 7 The image processing apparatus shown in the image processing method can be applied to virtual reality (VR) devices or augmented reality (AR) devices, and the model can be used to identify images to be monitored in virtual reality (VR) devices or augmented reality (AR) devices.
[0247] Figure 19 This is a schematic diagram of another image processing apparatus according to an embodiment of the present invention. Figure 19 As shown, the image processing device 1800 may include: a presentation unit 1902, a third recognition unit 1904, and a fifth training unit 1906.
[0248] The presentation unit 1902 is used to display an input image of the scene to be monitored on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device.
[0249] The third recognition unit 1904 is used to send input images to the edge device via virtual reality (VR) devices or augmented reality (AR) devices. The original cloud model at the edge device performs image recognition on the input image and trains the target cloud model based on the recognized feature vectors. The original cloud model is obtained by training the original edge model with a video dataset. The video dataset is obtained by the edge device through monitoring with different image acquisition devices at the first moment.
[0250] The fifth training unit 1906 is used to train the target cloud model on the video dataset monitored at the second time using different image acquisition devices, and after updating the original edge model based on the trained target edge model, drive the VR device or AR device to render and display the video dataset monitored at the third time.
[0251] It should be noted that the aforementioned presentation unit 1902, third recognition unit 1904, and fifth training unit 1906 correspond to steps S602 to S606 in Embodiment 1. The three units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the aforementioned units, as part of the device, can run in the computer terminal A provided in Embodiment 1.
[0252] In the image processing apparatus of this embodiment, the present invention obtains an original cloud model by training the original edge model using the video dataset monitored at the first moment without manual annotation, and obtains a target edge model by training the target cloud model using the video dataset monitored at the second moment. Thus, by using a large amount of video data to coordinate and update the original cloud model and the original edge model, the technical effect of timely model updating is achieved, solving the technical problem of low efficiency in model updating.
[0253] Example 4
[0254] Embodiments of the present invention may provide an image processor, which may include a computer terminal, which may be any computer terminal device from a group of computer terminals. Optionally, in this embodiment, the computer terminal may also be replaced by a mobile terminal or other terminal device.
[0255] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.
[0256] In this embodiment, the computer terminal described above can execute the program code for the following steps in the image processing method of the application: obtaining the original cloud model from the edge, wherein the original cloud model is obtained by training the original edge model with a video dataset, and the video dataset is obtained by the edge using different image acquisition devices to monitor at the first moment; performing image recognition on the input image in the scene to be monitored based on the original cloud model to obtain feature vectors; training the target cloud model based on the feature vectors; training the target cloud model with the video datasets monitored by different image acquisition devices at the second moment to obtain the target edge model, wherein the target edge model is used to update the original edge model, so that the edge can identify the video datasets monitored by different image acquisition devices at the third moment based on the target edge model.
[0257] Optionally, Figure 20 This is a structural block diagram of a computer terminal according to an embodiment of the present invention. Figure 20 As shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 2002, memory 2004, and transmission devices 2006.
[0258] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the image processing method and apparatus in this embodiment of the invention. The processor executes various functional applications and predictions by running the software programs and modules stored in the memory, thereby realizing the aforementioned image processing method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0259] As an alternative example, the processor can invoke information and applications stored in the memory via a transmission device to perform the following steps: acquiring the original cloud model from the edge, wherein the original cloud model is obtained by training the original edge model with a video dataset, the video dataset being obtained by the edge using different image acquisition devices at a first moment; performing image recognition on the input image in the scene to be monitored based on the original cloud model to obtain feature vectors; training a target cloud model based on the feature vectors; training the target cloud model with the video datasets monitored by different image acquisition devices at a second moment to obtain a target edge model, wherein the target edge model is used to update the original edge model, enabling the edge to recognize the video datasets monitored by different image acquisition devices at a third moment based on the target edge model.
[0260] Optionally, the processor may also execute program code that performs the following steps: based on the original cloud models corresponding to multiple monitoring scenarios, perform image recognition on the input images in the multiple monitoring scenarios to obtain multiple sub-feature vectors, wherein the multiple sub-feature vectors correspond one-to-one with the multiple monitoring scenarios; and perform average pooling on the multiple sub-feature vectors to obtain feature vectors.
[0261] Optionally, the processor may also execute program code that performs the following steps: training the original cloud model based on the feature vector and the target loss function to obtain the target cloud model.
[0262] Optionally, the processor may also execute program code for the following steps: identifying the monitored video dataset based on the original edge model to obtain the identification result; determining the identification result as a pseudo-label, and training the target cloud model based on the pseudo-label to obtain the target edge model; or performing transfer learning on the weights of the target cloud model based on the monitored video dataset to obtain the target edge model.
[0263] As an alternative example, the processor can invoke information and applications stored in the memory via a transmission device to perform the following steps: acquiring video datasets monitored by different image acquisition devices at a first moment; training the original edge model using the monitored video datasets to obtain an original cloud model, wherein the original cloud model is used to enable the cloud to train feature vectors of images based on input images in the scene to be monitored, thereby obtaining a target cloud model; acquiring the target edge model obtained by training the target cloud model using video datasets monitored by different image acquisition devices at a second moment; updating the original edge model to the target edge model, wherein the target edge model is used to identify the video datasets monitored by different image acquisition devices at a third moment.
[0264] Optionally, the processor may also execute program code that performs the following steps: determining the target edge model as the original edge model, determining the video dataset monitored at the third time as the video dataset monitored at the first time, and returning to execute the step of training the original edge model using the monitored video dataset to obtain the original cloud model.
[0265] Optionally, the processor may also execute program code for the following steps: augmenting the monitored video dataset to obtain an augmented video dataset; using the augmented video dataset as training samples to perform self-supervised training on the original edge model to obtain the original cloud model, wherein the distance between multiple positive samples in the training samples is less than the distance between multiple negative samples.
[0266] Optionally, the processor may also execute program code that performs the following steps: masking the monitored video dataset to obtain a masked video dataset; restoring the masked video dataset based on the original edge model; and training the original edge model based on the restored masked video dataset and the monitored video dataset to obtain the original cloud model.
[0267] Optionally, the processor may also execute program code that performs the following steps: identifies pseudo-labels of the monitored video dataset based on the original edge model; generates text from the pseudo-labels of the monitored video dataset, wherein the text is used to describe the monitored video dataset; and trains the original edge model into an original cloud model based on the monitored video dataset and the text.
[0268] As an alternative example, the processor can invoke information and applications stored in memory via a transmission device to perform the following steps: acquiring the original cloud model from the edge, wherein the original cloud model is obtained by training the original edge model with a video dataset, the video dataset being obtained by the edge using different image acquisition devices to monitor traffic roads at a first moment, and the video dataset containing at least one vehicle traveling through the traffic road; recognizing the input images in the traffic road based on the original cloud model to obtain feature vectors; training a target cloud model based on the feature vectors; training the target cloud model with the video dataset monitored by different image acquisition devices at a second moment to obtain a target edge model, wherein the target edge model is used to update the original edge model, so that the edge can recognize the video datasets monitored by different image acquisition devices on the traffic road at a third moment based on the target edge model.
[0269] As an alternative example, the processor can invoke information and applications stored in memory via a transmission device to perform the following steps: displaying an input image of the scene to be monitored on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device; the VR device or AR device sends the input image to the edge, wherein the edge's original cloud model performs image recognition on the input image and trains a target cloud model based on the recognized feature vectors, the original cloud model being trained on the original edge model using a video dataset, the video dataset being obtained by the edge using different image acquisition devices at a first moment; after training the target cloud model using the video dataset monitored by different image acquisition devices at a second moment, and updating the original edge model based on the trained target edge model, the processor drives the VR device or AR device to render and display the video dataset monitored at a third moment.
[0270] This invention, without requiring manual annotation, trains the original edge model using the video dataset monitored at the first moment to obtain the original cloud model, and trains the target cloud model using the video dataset monitored at the second moment to obtain the target edge model. By using a large amount of video data to coordinate and update the original cloud model and the original edge model, the technical effect of timely model updates is achieved, solving the technical problem of low efficiency in model updates.
[0271] Those skilled in the art will understand that Figure 20 The structure shown is for illustrative purposes only. Computer terminal A can also be a smartphone (such as an Android phone, iOS phone, etc.), tablet computer, mobile internet device (MID), PAD and other terminal devices. Figure 20This does not limit the structure of the aforementioned computer terminal A. For example, computer terminal A may also include components that are more complex than those described above. Figure 20 Showing more or fewer components (such as network interfaces, display devices, etc.), or having the same Figure 20 The different configurations shown.
[0272] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0273] Example 5
[0274] Embodiments of the present invention also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the image processing method provided in Embodiment 1.
[0275] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0276] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: obtaining an original cloud model from the edge, wherein the original cloud model is obtained by training an original edge model with a video dataset, the video dataset being obtained by the edge using different image acquisition devices at a first moment; performing image recognition on the input image in the scene to be monitored based on the original cloud model to obtain a feature vector; training a target cloud model based on the feature vector; training the target cloud model with the video dataset monitored by different image acquisition devices at a second moment to obtain a target edge model, wherein the target edge model is used to update the original edge model, so that the edge can identify the video dataset monitored by different image acquisition devices at a third moment based on the target edge model.
[0277] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: based on the original cloud models corresponding to multiple monitoring scenarios, image recognition is performed on the input images in the multiple monitoring scenarios to obtain multiple sub-feature vectors, wherein the multiple sub-feature vectors correspond one-to-one with the multiple monitoring scenarios; average pooling is performed on the multiple sub-feature vectors to obtain feature vectors.
[0278] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: training the original cloud model based on the feature vector and the target loss function to obtain the target cloud model.
[0279] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: identifying the monitored video dataset based on the original edge model to obtain identification results; determining the identification results as pseudo-labels and training the target cloud model based on the pseudo-labels to obtain the target edge model; or performing transfer learning on the weights of the target cloud model based on the monitored video dataset to obtain the target edge model.
[0280] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: acquiring a video dataset monitored by different image acquisition devices at a first moment; training an original edge model using the monitored video dataset to obtain an original cloud model, wherein the original cloud model is used to enable the cloud to train feature vectors of images based on input images in the scene to be monitored, thereby obtaining a target cloud model; acquiring a target edge model obtained by training the target cloud model using video datasets monitored by different image acquisition devices at a second moment; updating the original edge model to the target edge model, wherein the target edge model is used to identify the video datasets monitored by different image acquisition devices at a third moment.
[0281] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: determining the target edge model as the original edge model, determining the video dataset monitored at the third time as the video dataset monitored at the first time, and returning to execute the steps of training the original edge model using the monitored video dataset to obtain the original cloud model.
[0282] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: augmenting the monitored video dataset to obtain an augmented video dataset; using the augmented video dataset as training samples to perform self-supervised training on the original edge model to obtain the original cloud model, wherein the distance between multiple positive samples in the training samples is less than the distance between multiple negative samples.
[0283] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: masking the monitored video dataset to obtain a masked video dataset; restoring the masked video dataset based on the original edge model; and training the original edge model based on the restored masked video dataset and the monitored video dataset to obtain the original cloud model.
[0284] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: identifying pseudo-labels of the monitored video dataset based on the original edge model; generating text from the pseudo-labels of the monitored video dataset, wherein the text is used to describe the monitored video dataset; and training the original edge model into an original cloud model based on the monitored video dataset and the text.
[0285] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: obtaining an original cloud model from the edge, wherein the original cloud model is obtained by training an original edge model on a video dataset, the video dataset being obtained by the edge monitoring traffic roads at a first time using different image acquisition devices, the video dataset containing at least one vehicle traveling through the traffic roads; identifying input images in the traffic roads based on the original cloud model to obtain feature vectors; training a target cloud model based on the feature vectors; training the target cloud model on the target cloud model using video datasets monitored by different image acquisition devices at a second time to obtain a target edge model, wherein the target edge model is used to update the original edge model, such that the edge identifies video datasets monitored by different image acquisition devices on the traffic roads at a third time based on the target edge model.
[0286] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: displaying an input image of a scene to be monitored on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device; the VR device or AR device sending the input image to an edge, wherein the edge's original cloud model performs image recognition on the input image and trains a target cloud model based on the recognized feature vectors, the original cloud model being trained on the original edge model using a video dataset, the video dataset being obtained by the edge using different image acquisition devices at a first moment; after training the target cloud model using the video dataset monitored by different image acquisition devices at a second moment, and updating the original edge model based on the trained target edge model, driving the VR device or AR device to render and display the video data monitored at a third moment.
[0287] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0288] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0289] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0290] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0291] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0292] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0293] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. An image processing method, characterized in that, include: Obtain the original cloud model from the edge, wherein the original cloud model is obtained by training the original edge model with a video dataset, and the video dataset is obtained by the edge using different image acquisition devices to monitor at the first moment; Based on the original cloud model, image recognition is performed on the input image in the scene to be monitored to obtain feature vectors; The target cloud model is obtained by training based on the aforementioned feature vectors; The target cloud model is trained using the video datasets monitored by the different image acquisition devices at the second time to obtain the target edge model. The target edge model is used to update the original edge model, so that the edge can identify the video datasets monitored by the different image acquisition devices at the third time based on the target edge model. The step of performing image recognition on the input images in the monitored scene based on the original cloud model to obtain feature vectors includes: performing image recognition on the input images in the multiple monitored scenes based on the original cloud models corresponding to the multiple monitored scenes respectively, to obtain multiple sub-feature vectors, wherein the multiple sub-feature vectors correspond one-to-one with the multiple monitored scenes; and performing fusion processing on the multiple sub-feature vectors to obtain the feature vector.
2. The method according to claim 1, characterized in that, The process of fusing multiple sub-feature vectors to obtain the feature vector includes: The feature vector is obtained by performing average pooling on multiple sub-feature vectors.
3. The method according to claim 1, characterized in that, The target cloud model is trained based on the aforementioned feature vectors, including: The original cloud model is trained based on the feature vector and the target loss function to obtain the target cloud model.
4. The method according to claim 1, characterized in that, The target cloud model is trained using video datasets monitored by the different image acquisition devices at the second time point to obtain the target edge model, including: The monitored video dataset is identified based on the original edge model to obtain identification results; the identification results are determined as pseudo-labels, and the target cloud model is trained based on the pseudo-labels to obtain the target edge model; or The target cloud model is obtained by performing transfer learning on the weights of the target cloud model based on the monitored video dataset.
5. An image processing method, characterized in that, include: Acquire video datasets monitored in real time by different image acquisition devices; The original edge model is trained using the monitored video dataset to obtain the original cloud model. The original cloud model is used to train the feature vector of the cloud based on the input image in the scene to be monitored, thereby obtaining the target cloud model. The feature vector is obtained by fusing multiple sub-feature vectors. The multiple sub-feature vectors are obtained by image recognition of the input images in the multiple scenes to be monitored based on the original cloud model corresponding to the multiple scenes to be monitored, and the multiple sub-feature vectors correspond one-to-one with the multiple scenes to be monitored. The target edge model is obtained by training the target cloud model with video datasets monitored by the different image acquisition devices at the second moment in the cloud. The original edge model is updated to the target edge model, wherein the target edge model is used to identify the video datasets monitored by the different image acquisition devices at the third time.
6. The method according to claim 5, characterized in that, The original edge model is updated to the target edge model to identify the video datasets monitored by the different image acquisition devices at the third time moment, including: The target edge model is determined as the original edge model, and the video dataset monitored at the third time moment is determined as the video dataset monitored at the first time moment. Then, the process of training the original edge model using the monitored video dataset is returned to obtain the original cloud model.
7. The method according to claim 5, characterized in that, The original edge model is trained using the monitored video dataset to obtain the original cloud model, which includes one of the following: The monitored video dataset is augmented to obtain an augmented video dataset; the augmented video dataset is used as training samples to perform self-supervised training on the original edge model to obtain the original cloud model, wherein the distance between multiple positive samples in the training samples is smaller than the distance between multiple negative samples; The monitored video dataset is masked to obtain a masked video dataset; the masked video dataset is restored based on the original edge model; the original edge model is trained based on the restored masked video dataset and the monitored video dataset to obtain the original cloud model; Based on the original edge model, pseudo-labels of the monitored video dataset are identified; text is generated from the pseudo-labels of the monitored video dataset, wherein the text is used to describe the monitored video dataset; based on the monitored video dataset and the text, the original edge model is trained into the original cloud model.
8. An image processing system, characterized in that, include: The edge is used to acquire video datasets monitored by different image acquisition devices at the first moment, and the original edge model is trained using the monitored video datasets to obtain the original cloud model; The cloud is used to perform image recognition on the input image in the scene to be monitored based on the original cloud model to obtain feature vectors, train the target cloud model based on the feature vectors, and train the target cloud model using the video dataset monitored by the different image acquisition devices at the second time to obtain the target edge model, wherein the target edge model is used to update the original edge model; The edge is also used to identify the video datasets monitored by the different image acquisition devices at the third time based on the target edge model; The cloud platform performs image recognition on the input image in the scene to be monitored based on the original cloud platform model through the following steps to obtain the feature vector: based on the original cloud platform model corresponding to the multiple scenes to be monitored, image recognition is performed on the input images in the multiple scenes to be monitored to obtain multiple sub-feature vectors, wherein the multiple sub-feature vectors correspond one-to-one with the multiple scenes to be monitored; the multiple sub-feature vectors are fused to obtain the feature vector.
9. An image processing method, characterized in that, include: Obtain the original cloud model from the edge, wherein the original cloud model is obtained by training the original edge model with a video dataset, the video dataset is obtained by the edge using different image acquisition devices to monitor traffic roads at the first moment, and the video dataset contains at least one vehicle driving through the traffic road; Based on the original cloud model, image recognition is performed on the input images in the traffic road to obtain feature vectors; The target cloud model is obtained by training based on the aforementioned feature vectors; The target cloud model is trained using the video datasets monitored by the different image acquisition devices at the second time to obtain the target edge model. The target edge model is used to update the original edge model, so that the edge can identify the video datasets monitored by the different image acquisition devices for the traffic road at the third time based on the target edge model. The step of performing image recognition on the input images in the traffic roads based on the original cloud model to obtain feature vectors includes: performing image recognition on the input images in the multiple traffic roads based on the original cloud model corresponding to the multiple traffic roads respectively, to obtain multiple sub-feature vectors, wherein the multiple sub-feature vectors correspond one-to-one with the multiple traffic roads; and performing fusion processing on the multiple sub-feature vectors to obtain the feature vector.
10. An image processing method, characterized in that, include: Display the input image of the scene to be monitored on the screen of a virtual reality (VR) device or an augmented reality (AR) device; The virtual reality (VR) device or augmented reality (AR) device sends the input image to the edge, wherein the original cloud model of the edge performs image recognition on the input image and trains a target cloud model based on the recognized feature vectors. The original cloud model is obtained by training the original edge model with a video dataset. The video dataset is obtained by the edge using different image acquisition devices to monitor at a first moment. The feature vector is obtained by fusing multiple sub-feature vectors. The multiple sub-feature vectors are obtained by performing image recognition on the input images in the multiple scenes to be monitored based on the original cloud models corresponding to the multiple scenes to be monitored, and the multiple sub-feature vectors correspond one-to-one with the multiple scenes to be monitored. After training the target cloud model using the video datasets monitored by the different image acquisition devices at the second moment, and updating the original edge model based on the trained target edge model, the VR device or the AR device is driven to render and display the video dataset monitored at the third moment.