Data processing method and system, storage medium and mobile device

By adopting an image segmentation model based on a shared encoding network in the image cutout, the problem of long processing time in the prior art is solved, and a fast and real-time image cutout effect is achieved.

CN113763393BActive Publication Date: 2025-05-06ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010508056.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-05
Publication Date
2025-05-06
Estimated Expiration
2040-06-05

AI Technical Summary

Technical Problem

In the prior art, the network structure of the image cutting method is heavier and has high computational complexity, resulting in a long processing time, and it is impossible to quickly complete the cutting of the picture or realize real-time cutting.

Method used

The image segmentation model obtained by using the shared encoding network training is used to reduce the model size, reduce the calculation amount, and improve the forward performance of the model through the encoder-decoder model structure shared by the feature encoding module.

Benefits of technology

It realizes that while ensuring accuracy, reduce processing time, improve processing performance, and can quickly complete picture cutting or real-time picture cutting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113763393B_ABST
    Figure CN113763393B_ABST
Patent Text Reader

Abstract

The present application discloses a data processing method and system, a storage medium and a mobile device. The method comprises: receiving an image to be processed, wherein the image to be processed includes a target object; processing the image to be processed using an image segmentation model to determine a target area where the target object is located, wherein the image segmentation model is obtained based on a shared coding network training. The present application solves the technical problem that the processing flow of the image to be processed in the related art is highly complex, resulting in a long processing time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular, to a data processing method and system, a storage medium, and a mobile device. Background Art

[0002] With the development of technology and the promotion of the visual industry, image matting technology has been widely used, for example, in e-commerce scene mapping, photo editing plug-ins, terminal apps, mini-programs, etc. At the same time, users' requirements for effects have also changed, from framing the foreground in the image to segmenting the foreground and how to solve the problem of uncertain areas in the foreground (fused areas), and the requirements for the refinement of the visualization effect have become higher and higher.

[0003] Currently, image matting can be achieved through the following methods: semantic segmentation plus traditional matting method, end-to-end deep image matting method, etc. However, the network structure used in the above methods is relatively heavy and the computational complexity is high, resulting in a long processing time and unable to quickly complete the image matting or meet the requirements of real-time image matting.

[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0005] The embodiments of the present application provide a data processing method and system, a storage medium and a mobile device to at least solve the technical problem in the related art that the processing flow of the image to be processed is highly complex, resulting in a long processing time.

[0006] According to one aspect of an embodiment of the present application, a data processing method is provided, comprising: receiving a processing request; based on the processing request, obtaining an initial model and at least one set of training samples, wherein the training samples include: image data, and label information of an area where an object is located contained in the image data; based on a coding network shared in the initial model, training the initial model using at least one set of training samples to obtain an image segmentation model; and outputting the image segmentation model.

[0007] According to another aspect of an embodiment of the present application, a data processing method is also provided, including: receiving an image to be processed, wherein the image to be processed includes a target object; processing the image to be processed using an image segmentation model to determine a target area where the target object is located, wherein the image segmentation model is obtained based on training of a shared encoding network; and displaying an image of the target area.

[0008] According to another aspect of an embodiment of the present application, a data processing method is also provided, including: receiving an image to be processed, wherein the image to be processed includes a target object; obtaining a neural network model; inputting the image to be processed into a shared encoding network of the neural network model to obtain target features of the image to be processed; inputting the target features into a first decoding network of the neural network model to obtain a target area where the target object is located.

[0009] According to another aspect of an embodiment of the present application, a data processing method is also provided, including: obtaining an image to be processed, wherein the image to be processed includes a target object; processing the image to be processed using an image segmentation model to determine a target area where the target object is located, wherein the image segmentation model is obtained based on a shared encoding network training.

[0010] According to another aspect of an embodiment of the present application, a data processing method is also provided, including: obtaining at least one group of training samples, wherein the training samples include: image data, and label information of an area where an object is located contained in the image data; inputting at least one group of training samples into a shared encoding network of an image segmentation model to obtain feature data of at least one group of training samples; inputting the feature data into a first decoding network of the image segmentation model to determine the area where the object is located; inputting the feature data and output information of at least one feature layer in the first decoding network into a second decoding network to obtain a processing result of the area; and updating the network weights of the image segmentation model based on the processing result of the area and the label information of the area where the object is located.

[0011] According to another aspect of an embodiment of the present application, a data processing method is also provided, including: receiving data to be processed; obtaining a neural network model; inputting the data to be processed into a shared encoding network of the neural network model to obtain target features of the data to be processed; inputting the target features into a first decoding network of the neural network model to obtain a processing result.

[0012] According to another aspect of an embodiment of the present application, a storage medium is further provided, the storage medium including a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the above-mentioned data processing method.

[0013] According to another aspect of an embodiment of the present application, a mobile device is provided, including a memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the above-mentioned data processing method is executed when the program is run.

[0014] According to another aspect of an embodiment of the present application, a data processing system is also provided, including: a processor; and a memory connected to the processor, for providing the processor with instructions for processing the following processing steps: obtaining an image to be processed, wherein the image to be processed includes a target object; processing the image to be processed using an image segmentation model to determine a target area where the target object is located, wherein the image segmentation model is obtained based on training of a shared encoding network.

[0015] In an embodiment of the present application, after acquiring the image to be processed, the image to be processed can be processed using an image segmentation model to determine the target area where the target object is located, thereby achieving the purpose of image cutout on the terminal.

[0016] It is easy to notice that the image segmentation model is trained based on a shared coding network. Compared with the existing technology, the training accuracy of the image segmentation model can be improved through the shared coding network, and in the process of processing the image to be processed, the network time consumption can be saved, thereby achieving the technical effect of saving processing time and improving processing performance, and realizing the needs of quickly completing cutouts or real-time cutouts.

[0017] The solution provided by the embodiment of the present application solves the technical problem in the related art that the processing flow of the image to be processed is highly complex, resulting in a long processing time. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0019] Figure 1 It is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method according to an embodiment of the present application;

[0020] Figure 2 is a flow chart of a data processing method according to an embodiment of the present application;

[0021] Figure 3a is a schematic diagram of an optional image to be processed according to an embodiment of the present application;

[0022] Figure 3b is a schematic diagram of an optional target area according to an embodiment of the present application;

[0023] Figure 3c is a schematic diagram of an optional target area positioning according to an embodiment of the present application;

[0024] Figure 3d A schematic diagram of an optional target region regression result according to an embodiment of the present application;

[0025] Figure 4 is a schematic diagram of an optional data processing method according to an embodiment of the present application;

[0026] Figure 5 is a flow chart of another data processing method according to an embodiment of the present application;

[0027] Figure 6 is a schematic diagram of an optional interactive interface according to an embodiment of the present application;

[0028] Figure 7 is a flow chart of another data processing method according to an embodiment of the present application;

[0029] Figure 8 is a schematic diagram of a data processing device according to an embodiment of the present application;

[0030] Fig. 9 is a schematic diagram of another data processing device according to an embodiment of the present application;

[0031] Fig.10 is a schematic diagram of another data processing device according to an embodiment of the present application;

[0032] Fig.11 is a flowchart of a fourth data processing method according to an embodiment of the present application;

[0033] Fig.12 is a schematic diagram of an optional data processing method according to an embodiment of the present application;

[0034] Fig.13 is a flowchart of a fifth data processing method according to an embodiment of the present application;

[0035] Fig.14 is a flowchart of a sixth data processing method according to an embodiment of the present application; and

[0036] Fig.15 It is a structural block diagram of a computer terminal according to an embodiment of the present application. DETAILED DESCRIPTION

[0037] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0038] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0039] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:

[0040] Cutout: It can be to separate a part of an image or video from the original image or video into a separate layer.

[0041] Encoder-decoder model: It consists of an encoder and a decoder. The encoder is used to encode the input image to obtain a vector with semantics; the decoder is used to take the semantic vector generated by the encoder as input and output the target image.

[0042] Example 1

[0043] Currently, most users use adobe-ps or cloud plug-ins to cut out images, which is very time-consuming. Even a high-performance API takes nearly 1 second to cut out images. The demand for high-performance cutouts is increasing. Moreover, with the development of smart and mobile devices, it has become a trend to use algorithms to calculate CPU power on the end.

[0044] However, the network structure used in the cutout method in the related art is relatively heavy and has high computational complexity, resulting in a long processing time and being unable to quickly complete the cutout or meet the requirements of real-time cutout.

[0045] In order to solve the above problems, the present application provides an encoder-decoder model that shares a feature encoding module, which can reduce the model size and thus reduce the amount of calculation and improve the forward performance of the model while ensuring that the accuracy does not drop seriously, that is, the accuracy meets user requirements.

[0046] According to an embodiment of the present application, a data processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0047] The method embodiments provided in the embodiments of the present application can be executed in a mobile device, a computer terminal or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0048] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0049] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the data processing method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the above-mentioned data processing method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0050] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0051] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0052] It should be noted that, in some optional embodiments, the above Figure 1 The computer device (or mobile device) shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. It should be noted that Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the above-described computer device (or mobile device).

[0053] Under the above operating environment, this application provides Figure 2 The data processing method shown. Figure 2 is a flow chart of a data processing method according to an embodiment of the present application. Figure 2 As shown, the method comprises the following steps:

[0054] Step S22, obtaining an image to be processed, wherein the image to be processed includes a target object.

[0055] The image to be processed in the above steps may be an image to be cut out, which may be an image captured by a user through a mobile device or a robot such as a smart speaker, or an image searched by a user on a mobile device, but is not limited thereto. The target object in the above steps may be an object to be cut out, for example, a person, an animal, or other objects, but is not limited thereto. For example, Figure 3a In the image to be processed shown, the target object may be a person in the image.

[0056] For example, in a live broadcast of goods, users can take pictures of the goods sold by the anchor during the live broadcast, obtain an image containing the goods, and obtain the area where the goods are located by cutting out the image. For another example, in a smart speaker scenario, users can control the smart speaker through voice to take an image of the user, and obtain the area where the user is located by cutting out the image. Furthermore, the purpose of taking ID photos can be achieved by changing the background color.

[0057] Step S24, using the image segmentation model to process the image to be processed to determine the target area where the target object is located, wherein the image segmentation model is obtained by training based on a shared encoding network.

[0058] The target area in the above step can be the contour area of ​​the target object, rather than just an area that can frame the target object, for example Figure 3b-3d As shown, it is possible to avoid the cutout image from including background images or images of other objects other than the target object.

[0059] For mobile devices, the network structure of the image segmentation model is required to be smaller, for example, it is generally required to be within 5M. Therefore, in the embodiment of the present application, the image segmentation model can adopt an encoder-2decoder model structure shared by a feature encoding model, wherein the encoder model is shared and two decoder models are connected externally, namely decoder-one and decoder-two, decoder-one is used for forward prediction of the target area, and decoder-two is used to complete accurate regional regression.

[0060] In an optional embodiment, in order to ensure the processing accuracy of the image segmentation model and save processing time, the encoder model structure can be shared during the training process, and the mask accuracy of the target area prediction can be improved through two decoder models, thereby improving the accuracy and stability of the hyperparameters; and in the prediction process, it is implemented through the encoder model and the decoder-one model, and there is no need to introduce the decoder-two model, thereby improving the computing performance.

[0061] For example, Figure 4The image segmentation model shown in the figure is used as an example to illustrate that the encoder model is shared and two decoder models are connected. The decoder-one model completes the overall semantic matting process, and the decoder-two model completes the regional matting process. When predicting forward, the overall image prediction can be performed directly to reduce the amount of calculation. During the training process, the encoder model + decoder-one model can be used to complete the initial position of the target area and the initial semantic segmentation information, and complete the positioning of the target area of ​​the entire image. Figure 3c As shown in the box. Then the features in the encoder model + decoder-one model are shared with the decoder-two model for accurate regional regression. Therefore, the decoder-two model is beneficial to the regional accuracy information learning of the encoder model + decoder-one model. In the actual processing process, the encoder model + decoder-one model is used to complete it.

[0062] The solution provided in the above-mentioned embodiment 1 of the present application can use the image segmentation model to process the image to be processed after acquiring the image to be processed, and determine the target area where the target object is located, thereby achieving the purpose of image cutout on the terminal.

[0063] It is easy to notice that the image segmentation model is trained based on a shared coding network. Compared with the existing technology, the training accuracy of the image segmentation model can be improved through the shared coding network, and in the process of processing the image to be processed, the network time consumption can be saved, thereby achieving the technical effect of saving processing time and improving processing performance, and realizing the needs of quickly completing cutouts or real-time cutouts.

[0064] Therefore, the solution of the above-mentioned embodiment 1 provided in the present application solves the technical problem in the related art that the processing flow of the image to be processed is highly complex, resulting in a long processing time.

[0065] In the above embodiment of the present application, the image to be processed is processed using an image segmentation model to determine the target area where the target object is located, including: inputting the image to be processed into a shared encoding network to obtain the target features of the image to be processed; inputting the target features into a first decoding network to obtain the target area.

[0066] The encoding network in the above steps can be as follows Figure 4 The encoder model shown in Figure 2 can be a first decoding network. Figure 4 The decoder-one model shown.

[0067] For example, Figure 4The image segmentation model shown in FIG. 1 is used as an example to illustrate that after obtaining the image to be processed, Figure 3a As shown in , the image to be processed can be input into the encoder model to obtain the target features, and then the target features can be input into the decoder-one model to obtain the target area where the target object is located, as well as the semantic segmentation information, and complete the positioning project of the target object in the entire image, as shown in Figure 3b shown.

[0068] In the above embodiment of the present application, the method also includes: obtaining at least one group of training samples, wherein the training samples include: image data, and label information of the area where the object is located contained in the image data; inputting at least one group of training samples into a shared encoding network to obtain feature data of at least one group of training samples; inputting the feature data into a first decoding network to determine the area where the object is located; inputting the feature data and output information of at least one feature layer in the first decoding network into a second decoding network to obtain a processing result of the area; and updating the network weights of the image segmentation model based on the processing result of the area and the label information of the area where the object is located.

[0069] The label information in the above step may be the result of manually marking the region. The second decoding network in the above step may be as follows Figure 4 The decoder-two model shown is used to improve the training accuracy of the target area. The input of the model can be the features of the target area in the original image and the original image, and participates in the network weight update during the training process.

[0070] For example, Figure 4 The image segmentation model shown in FIG. 1 is used as an example to illustrate that in the training process of the image segmentation model, the training samples can be input into the encoder model, and the obtained feature data is first input into the decoder-one model, and then the feature layer information in the encoder model + decoder-one model is shared with the decoder-two model to obtain accurate regional regression, that is, to obtain the processing result of the region, for example, Figure 3d As shown, the network weights of the entire image segmentation model are further updated based on the processing results of the region to achieve the purpose of model training.

[0071] In the above embodiment of the present application, updating the network weights of the image segmentation model based on the processing results of the region and the label information of the region where the object is located includes: obtaining the loss value of the image segmentation model based on the processing results of the region and the label information of the region where the object is located; when the loss value is greater than a preset threshold, updating the network weights of the image segmentation model; when the loss value is less than the preset threshold, stopping updating the network weights of the image segmentation model.

[0072] The preset threshold in the above steps can be set in advance according to the processing accuracy of the image cutout, and the higher the accuracy, the smaller the preset threshold. In order to improve the forward performance of the model, the model processing accuracy can be reduced according to the actual situation.

[0073] In an optional embodiment, the overlap and similarity of two regions can be calculated based on the prediction results of the regions and the results of manual annotation to obtain the loss value of the image segmentation model. The loss value can be further compared with a preset threshold to determine whether to train the image segmentation model. If the loss value is greater than the preset threshold, it indicates that the processing accuracy of the image segmentation model cannot meet the requirements and the image segmentation model needs to continue to be trained. If the loss value is less than or equal to the preset threshold, it indicates that the processing accuracy of the image segmentation model can meet the requirements and the image segmentation model training is completed.

[0074] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0075] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0076] Example 2

[0077] According to an embodiment of the present application, a data processing method is also provided.

[0078] Figure 5 FIG. 1 is a flow chart of another data processing method according to an embodiment of the present application. Figure 5 As shown, the method comprises the following steps:

[0079] Step S52: receiving an image to be processed, wherein the image to be processed includes a target object.

[0080] The image to be processed in the above steps may be an image to be cut out, which may be an image captured by a user through a mobile device or a robot such as a smart speaker, or an image searched by a user on a mobile device, but is not limited thereto. The target object in the above steps may be an object to be cut out, for example, a person, an animal, or other objects, but is not limited thereto. For example, Figure 3a In the image to be processed shown, the target object may be a person in the image.

[0081] In an optional embodiment, in order to facilitate the user to cut out the picture, the mobile device (such as an Android phone, an IOS phone, a tablet computer, a laptop computer, etc., but not limited to this) can use an application, a small program, etc. to achieve the purpose of cutting out the picture. On this basis, the mobile device can provide the user with an interactive interface, such as Figure 6 As shown, the user can click the "shoot" button in the interactive interface to call the camera to shoot an image as the image to be processed; the user can also click the "open file" button in the interactive interface to query the storage space of the mobile device and select an image as the image to be processed.

[0082] Step S54, using the image segmentation model to process the image to be processed to determine the target area where the target object is located, wherein the image segmentation model is obtained by training based on a shared encoding network.

[0083] The target area in the above step can be the contour area of ​​the target object, rather than just an area that can frame the target object, for example Figure 3b-3d As shown, it is possible to avoid the cutout image from including background images or images of other objects other than the target object.

[0084] Step S56, displaying the image of the target area.

[0085] In an optional embodiment, if Figure 6 As shown, in order to facilitate users to view the cutout results, the mobile device can display the cutout image, that is, the image of the target area, in the output area of ​​the interactive interface, and the user can adjust the image displayed in the output area as needed to obtain the final cutout result.

[0086] The solution provided in the above-mentioned embodiment 2 of the present application can, after acquiring the image to be processed, use the image segmentation model to process the image to be processed, determine the target area where the target object is located, and display the image of the target area, thereby achieving the purpose of image cutout on the terminal.

[0087] It is easy to notice that the image segmentation model is trained based on a shared coding network. Compared with the existing technology, the training accuracy of the image segmentation model can be improved through the shared coding network, and in the process of processing the image to be processed, the network time consumption can be saved, thereby achieving the technical effect of saving processing time and improving processing performance, and realizing the needs of quickly completing cutouts or real-time cutouts.

[0088] Therefore, the solution of the above-mentioned embodiment 2 provided in the present application solves the technical problem in the related art that the processing flow of the image to be processed is highly complex, resulting in a long processing time.

[0089] In the above embodiment of the present application, after receiving the image to be processed, the method also includes: outputting an image segmentation model; receiving a selected target decoding network, wherein the target decoding network is the first decoding network or the second decoding network in the image segmentation model; inputting the image to be processed into the shared encoding network to obtain target features of the image to be processed; and processing the target features using the target decoding network to obtain a target area.

[0090] The encoding network in the above steps can be as follows Figure 4 The encoder model shown in Figure 2 can be a first decoding network. Figure 4 The decoder-one model shown in FIG. 1 is a decoder-one model. The second decoding network can be as follows: Figure 4 The decoder-two model shown is used to improve the training accuracy of the target area. The input of the model can be the features of the target area in the original image and the original image, and participates in the network weight update during the training process.

[0091] It should be noted that different users often have different requirements for the time and accuracy of cutouts. For example, some users hope to cutout images quickly but do not require high accuracy, while some users have high requirements for cutout accuracy.

[0092] In an optional embodiment, in order to meet the needs of different users, the image segmentation model used for cutout can be displayed to the user, and the user is informed that the model contains two decoding networks, and the two decoding networks correspond to different cutout accuracies, so that the user can choose according to his own needs. After the user has made the selection, the decoding network selected by the user can be used as the target decoding network, and the target features output by the encoding network are processed by the target decoding network to obtain the target area where the target object is located.

[0093] In the above embodiment of the present application, the target feature is processed by using the target decoding network to obtain the target area, including: when the target decoding network is a first decoding network, the target feature is input into the first decoding network to obtain the target area; when the target decoding network is a second decoding network, the target feature is input into the first decoding network, and the target feature and output information of at least one feature layer in the first decoding network are input into the second decoding network to obtain the target area.

[0094] In an optional embodiment, when the user selects the decoder-one model, it indicates that the user needs to quickly cut out the image, and the target features can be directly input into the decoder-one model to obtain the target area where the target object is located, as well as the semantic segmentation information, and complete the positioning project of the target object of the entire image. At this time, there is no need to use the decoder-two model, and the processing speed is faster. In another optional embodiment, when the user selects the decoder-two model, it indicates that the user needs high-precision cutouts, and the target features are first input into the decoder-one model, and then the feature layer information in the encoder model + decoder-one model is shared with the decoder-two model to obtain the final processing result.

[0095] In the above embodiment of the present application, the method also includes: obtaining at least one group of training samples, wherein the training samples include: image data, and label information of the area where the object is located contained in the image data; inputting at least one group of training samples into a shared encoding network to obtain feature data of at least one group of training samples; inputting the feature data into a first decoding network to determine the area where the object is located; inputting the feature data and output information of at least one feature layer in the first decoding network into a second decoding network to obtain a processing result of the area; and updating the network weights of the image segmentation model based on the processing result of the area and the label information of the area where the object is located.

[0096] In the above embodiment of the present application, updating the network weights of the image segmentation model based on the processing results of the region and the label information of the region where the object is located includes: obtaining the loss value of the image segmentation model based on the processing results of the region and the label information of the region where the object is located; when the loss value is greater than a preset threshold, updating the network weights of the image segmentation model; when the loss value is less than the preset threshold, stopping updating the network weights of the image segmentation model.

[0097] The preset threshold in the above steps can be set in advance according to the processing accuracy of the image cutout, and the higher the accuracy, the smaller the preset threshold. In order to improve the forward performance of the model, the model processing accuracy can be reduced according to the actual situation.

[0098] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0099] Example 3

[0100] According to an embodiment of the present application, a data processing method is also provided.

[0101] Figure 7 FIG. 1 is a flow chart of another data processing method according to an embodiment of the present application. Figure 7 As shown, the method comprises the following steps:

[0102] Step S72: obtaining at least one set of training samples, wherein the training samples include: image data, and label information of the area where the object is located contained in the image data.

[0103] The label information in the above step may be the result of manually marking the area.

[0104] Step S74: input at least one set of training samples into a shared encoding network of the image segmentation model to obtain feature data of at least one set of training samples.

[0105] The encoding network in the above steps can be as follows Figure 4 The encoder model shown.

[0106] Step S76, inputting the feature data into the first decoding network of the image segmentation model to determine the area where the object is located.

[0107] The first decoding network in the above steps can be as follows Figure 4 The decoder-one model shown in FIG. 1 is a decoder-one model. The region in the above step can be the contour region of the object, not just a region that can frame the target object, for example, Figure 3b-3d As shown, it is possible to avoid the cutout image from including background images or images of other objects other than the target object.

[0108] Step S78, input the feature data and output information of at least one feature layer in the first decoding network into the second decoding network to obtain a processing result of the region.

[0109] The second decoding network in the above steps can be as follows Figure 4 The decoder-two model shown is used to improve the training accuracy of the target area. The input of the model can be the features of the target area in the original image and the original image, and participates in the network weight update during the training process.

[0110] Step S710, based on the processing result of the region and the label information of the region where the object is located, the network weights of the image segmentation model are updated.

[0111] The solution provided in the above-mentioned embodiment 3 of the present application can, after acquiring at least one group of training samples, input at least one group of training samples into the shared encoding network of the image segmentation model to obtain feature data of at least one group of training samples, input the feature data into the first decoding network of the image segmentation model to determine the area where the object is located, input the feature data and the output information of at least one feature layer in the first decoding network into the second decoding network to obtain the processing result of the area, and based on the processing result of the area and the label information of the area where the object is located, update the network weights of the image segmentation model, thereby achieving the purpose of image matting on the terminal.

[0112] It is easy to notice that the training process of the image segmentation model can be realized through a shared encoding network, a first decoding network and a second decoding network. Compared with the prior art, the training accuracy of the image segmentation model can be improved through a shared encoding network, and in the process of processing the image to be processed, the network time consumption can be saved, thereby achieving the technical effect of saving processing time and improving processing performance, and realizing the needs of quickly completing cutouts or real-time cutouts.

[0113] Therefore, the solution of the above-mentioned embodiment 3 provided in the present application solves the technical problem in the related art that the processing flow of the image to be processed is highly complex, resulting in a long processing time.

[0114] In the above embodiment of the present application, updating the network weights of the image segmentation model based on the processing results of the region and the label information of the region where the object is located includes: obtaining the loss value of the image segmentation model based on the processing results of the region and the label information of the region where the object is located; when the loss value is greater than a preset threshold, updating the network weights of the image segmentation model; when the loss value is less than the preset threshold, stopping updating the network weights of the image segmentation model.

[0115] The preset threshold in the above steps can be set in advance according to the processing accuracy of the image cutout, and the higher the accuracy, the smaller the preset threshold. In order to improve the forward performance of the model, the model processing accuracy can be reduced according to the actual situation.

[0116] In the above embodiment of the present application, after the network weights of the image segmentation model are updated, the method also includes: obtaining an image to be processed, wherein the image to be processed includes a target object; inputting the image to be processed into a shared encoding network to obtain target features of the image to be processed; inputting the target features into a first decoding network to obtain a target area where the target object is located.

[0117] The image to be processed in the above steps may be an image to be cut out, which may be an image captured by a user through a mobile device or a robot such as a smart speaker, or an image searched by a user on a mobile device, but is not limited thereto. The target object in the above steps may be an object to be cut out, for example, a person, an animal, or other objects, but is not limited thereto. For example, Figure 3a In the image to be processed shown, the target object may be a person in the image.

[0118] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0119] Example 4

[0120] According to an embodiment of the present application, a data processing device for implementing the above data processing method is also provided. Figure 8 As shown, the device 800 includes: a first acquisition module 802 and a processing module 804.

[0121] Among them, the first acquisition module 802 is used to acquire the image to be processed, wherein the image to be processed includes the target object; the processing module 804 is used to process the image to be processed using an image segmentation model to determine the target area where the target object is located, wherein the image segmentation model is obtained based on shared encoding network training.

[0122] It should be noted that the first acquisition module 802 and the processing module 804 correspond to steps S22 to S24 in Example 1, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.

[0123] In the above embodiment of the present application, the processing module includes: a first input unit, used to input the image to be processed into a shared encoding network to obtain the target features of the image to be processed; and a second input unit, used to input the target features into a first decoding network to obtain the target area.

[0124] In the above embodiment of the present application, the device also includes: a second acquisition module, used to acquire at least one group of training samples, wherein the training samples include: image data, and label information of the area where the object is located contained in the image data; a first input module, used to input at least one group of training samples into a shared encoding network to obtain feature data of at least one group of training samples; a second input module, used to input the feature data into a first decoding network to determine the area where the object is located; a third input module, used to input the feature data and the output information of at least one feature layer in the first decoding network into the second decoding network to obtain the processing result of the area; an update module, used to update the network weights of the image segmentation model based on the processing result of the area and the label information of the area where the object is located.

[0125] In the above embodiment of the present application, the update module includes: a processing unit, which is used to obtain the loss value of the image segmentation model based on the processing result of the region and the label information of the region where the object is located; an update unit, which is used to update the network weights of the image segmentation model when the loss value is greater than a preset threshold; and a stop unit, which is used to stop updating the network weights of the image segmentation model when the loss value is less than a preset threshold.

[0126] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0127] Example 5

[0128] According to an embodiment of the present application, a data processing device for implementing the above data processing method is also provided. Fig. 9 As shown, the device 900 includes: a receiving module 902 , a processing module 904 and a display module 906 .

[0129] Among them, the receiving module 902 is used to receive the image to be processed, wherein the image to be processed includes the target object; the processing module 904 is used to process the image to be processed using an image segmentation model to determine the target area where the target object is located, wherein the image segmentation model is obtained based on a shared encoding network training; the display module 906 is used to display the image of the target area.

[0130] It should be noted that the receiving module 902, the processing module 904 and the display module 906 correspond to steps S52 to S56 in Example 2, and the three modules and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.

[0131] In the above embodiment of the present application, the processing module includes: a first input unit, used to input the image to be processed into a shared encoding network to obtain the target features of the image to be processed; and a second input unit, used to input the target features into a first decoding network to obtain the target area.

[0132] In the above embodiment of the present application, the device also includes: an acquisition module, used to acquire at least one group of training samples, wherein the training samples include: image data, and label information of the area where the object is located contained in the image data; a first input module, used to input at least one group of training samples into a shared encoding network to obtain feature data of at least one group of training samples; a second input module, used to input the feature data into a first decoding network to determine the area where the object is located; a third input module, used to input the feature data and the output information of at least one feature layer in the first decoding network into the second decoding network to obtain the processing result of the area; an update module, used to update the network weights of the image segmentation model based on the processing result of the area and the label information of the area where the object is located.

[0133] In the above embodiment of the present application, the update module includes: a processing unit, which is used to obtain the loss value of the image segmentation model based on the processing result of the region and the label information of the region where the object is located; an update unit, which is used to update the network weights of the image segmentation model when the loss value is greater than a preset threshold; and a stop unit, which is used to stop updating the network weights of the image segmentation model when the loss value is less than a preset threshold.

[0134] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0135] Example 6

[0136] According to an embodiment of the present application, a data processing device for implementing the above data processing method is also provided. Fig.10 As shown, the device 1000 includes: a first acquisition module 1002 , a first input module 1004 , a second input module 1006 , a third input module 1008 and an update module 1010 .

[0137] Among them, the first acquisition module 1002 is used to obtain at least one group of training samples, wherein the training samples include: image data, and label information of the area where the object is located contained in the image data; the first input module 1004 is used to input at least one group of training samples into the shared encoding network to obtain feature data of at least one group of training samples; the second input module 1006 is used to input the feature data into the first decoding network to determine the area where the object is located; the third input module 1008 is used to input the feature data and the output information of at least one feature layer in the first decoding network into the second decoding network to obtain the processing result of the area; the update module 1010 is used to update the network weights of the image segmentation model based on the processing result of the area and the label information of the area where the object is located.

[0138] It should be noted that the first acquisition module 1002, the first input module 1004, the second input module 1006, the third input module 1008 and the update module 1010 correspond to steps S72 to S710 in Example 3, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules as part of the device can be run in the computer terminal 10 provided in Example 1.

[0139] In the above embodiment of the present application, the update module includes: a processing unit, which is used to obtain the loss value of the image segmentation model based on the processing result of the region and the label information of the region where the object is located; an update unit, which is used to update the network weights of the image segmentation model when the loss value is greater than a preset threshold; and a stop unit, which is used to stop updating the network weights of the image segmentation model when the loss value is less than a preset threshold.

[0140] In the above embodiment of the present application, the device also includes: a second acquisition module, used to acquire the image to be processed, wherein the image to be processed includes the target object; a fourth input module, used to input the image to be processed into the shared encoding network to obtain the target features of the image to be processed; and a fifth input module, used to input the target features into the first decoding network to obtain the target area where the target object is located.

[0141] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0142] Example 7

[0143] According to an embodiment of the present application, a data processing system is also provided, including:

[0144] Processor. And

[0145] The memory is connected to the processor and is used to provide the processor with instructions for processing the following processing steps: obtaining an image to be processed, wherein the image to be processed includes a target object; using an image segmentation model to process the image to be processed to determine a target area where the target object is located, wherein the image segmentation model is obtained based on a shared encoding network training.

[0146] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0147] Example 8

[0148] According to an embodiment of the present application, a data processing method is also provided.

[0149] Fig.11 is a flow chart of a fourth data processing method according to an embodiment of the present application. Fig.11 As shown, the method comprises the following steps:

[0150] Step S112, receiving a processing request.

[0151] The processing request in the above steps may be a request for building a network model, which may carry the data to be processed and the corresponding processing results, etc., or a request for training a network model, which may carry a constructed initial model and the specific role of the initial model, etc. In the embodiment of the present application, the construction of an image segmentation model is taken as an example for explanation.

[0152] In an optional embodiment, for ordinary users or image cutting software companies, due to the limitation on the number of images, in order to improve data processing accuracy, a model interface can be provided to users, and users upload requests to build network models to the server through the interface, so that the server can build and train network models for users according to the user's request.

[0153] Step S114, based on the processing request, obtaining an initial model and at least one set of training samples, wherein the training samples include: image data, and label information of the area where the object is located contained in the image data.

[0154] The label information in the above steps can be the result of manually annotating the region. The above initial model can be used as follows Fig.12The network structure shown is an encoder-2decoder model structure that uses a shared feature encoding model, in which the encoder model is shared and two decoder models are connected, namely decoder-one and decoder-two. Decoder-one is used for forward prediction of the target area, while decoder-two is used to complete accurate regional regression.

[0155] Step S116, based on the encoding network shared in the initial model, the initial model is trained using at least one group of training samples to obtain an image segmentation model.

[0156] The encoding network in the above steps can be as follows Fig.12 The encoder model shown.

[0157] For example, Fig.12 The image segmentation model shown in the figure is used as an example to illustrate that the encoder model is shared and two decoder models are connected. The decoder-one model completes the overall semantic matting process, and the decoder-two model completes the regional matting process. During the training process, the encoder model + decoder-one model can be used to complete the initial position of the target area and the initial semantic segmentation information, and complete the positioning project of the target area of ​​the entire image, such as Figure 3c As shown in the box, the features in the encoder model + decoder-one model are then shared with the decoder-two model for accurate regional regression.

[0158] Step S118, output the image segmentation model.

[0159] In an optional embodiment, after the initial model training is completed, the trained image segmentation model can be returned to the user, and the user can perform operations such as cutout by himself.

[0160] In the above embodiment of the present application, based on the shared encoding network in the initial model, the initial model is trained using at least one group of training samples to obtain an image segmentation model, including: inputting at least one group of training samples into the shared encoding network to obtain feature data of at least one group of training samples; inputting the feature data into the first decoding network of the initial model to determine the area where the object is located; inputting the feature data and the output information of at least one feature layer in the first decoding network into the second decoding network of the initial model to obtain the processing result of the area; based on the processing result of the area and the label information of the area where the object is located, the network weights of the initial model are updated to obtain the image segmentation model.

[0161] The first decoding network in the above steps can be as follows Fig.12 In the decoder-one model shown in FIG. 1 , the region in the above step can be the contour region of the object, rather than just an area that can frame the target object. The second decoding network can be as follows: Fig.12 The decoder-two model shown is used to improve the training accuracy of the target area. The input of the model can be the features of the target area in the original image and the original image, and participates in the network weight update during the training process.

[0162] For example, Fig.12 The image segmentation model shown in FIG. 1 is used as an example to illustrate that during the training process, the training samples can be input into the encoder model, and the obtained feature data is first input into the decoder-one model, and then the feature layer information in the encoder model + decoder-one model is shared with the decoder-two model to obtain accurate regional regression, that is, to obtain the processing result of the region, for example, Figure 3d As shown, the network weights of the entire image segmentation model are further updated based on the processing results of the region to achieve the purpose of model training.

[0163] In the above embodiment of the present application, updating the network weights of the initial model based on the processing results of the region and the label information of the region where the object is located includes: obtaining the loss value of the initial model based on the processing results of the region and the label information of the region where the object is located; when the loss value is greater than a preset threshold, updating the network weights of the initial model; when the loss value is less than the preset threshold, stopping updating the network weights of the initial model.

[0164] The preset threshold in the above steps can be set in advance according to the processing accuracy of the image cutout, and the higher the accuracy, the smaller the preset threshold. In order to improve the forward performance of the model, the model processing accuracy can be reduced according to the actual situation.

[0165] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0166] Example 9

[0167] According to an embodiment of the present application, a data processing method is also provided.

[0168] Fig.13 is a flow chart of a fifth data processing method according to an embodiment of the present application. Fig.13 As shown, the method comprises the following steps:

[0169] Step S132: receiving an image to be processed, wherein the image to be processed includes a target object.

[0170] The image to be processed in the above steps may be an image to be cut out, which may be an image captured by a user through a mobile device or a robot such as a smart speaker, or an image searched by a user on a mobile device, but is not limited thereto. The target object in the above steps may be an object to be cut out, for example, a person, an animal, or other objects, but is not limited thereto. For example, Figure 3a In the image to be processed shown, the target object may be a person in the image.

[0171] Step S134, obtaining a neural network model.

[0172] The neural network model in the above steps can adopt an encoder-2decoder model structure shared by a feature encoding model, wherein the encoder model is shared and two decoder models are connected externally, namely decoder-one and decoder-two. Decoder-one is used for forward prediction of the target area, while decoder-two is used to complete accurate regional regression.

[0173] Step S136, inputting the image to be processed into the shared encoding network of the neural network model to obtain the target features of the image to be processed.

[0174] The encoding network in the above steps can be as follows Figure 4 The encoder model shown.

[0175] Step S138, inputting the target features into the first decoding network of the neural network model to obtain the target area where the target object is located.

[0176] The first decoding network in the above steps can be as follows Figure 4 The target region can be the contour region of the target object, rather than just a region that can frame the target object, for example, Figure 3b-3d As shown, it is possible to avoid the cutout image from including background images or images of other objects other than the target object.

[0177] In the above embodiment of the present application, the method also includes: obtaining at least one group of training samples, wherein the training samples include: image data, and label information of the area where the object is located contained in the image data; inputting at least one group of training samples into a shared encoding network to obtain feature data of at least one group of training samples; inputting the feature data into a first decoding network to determine the area where the object is located; inputting the feature data and output information of at least one feature layer in the first decoding network into a second decoding network of a neural network model to obtain a processing result of the area; and updating the network weights of the neural network model based on the processing result of the area and the label information of the area where the object is located.

[0178] In the above embodiment of the present application, updating the network weights of the neural network model based on the processing results of the area and the label information of the area where the object is located includes: obtaining the loss value of the neural network model based on the processing results of the area and the label information of the area where the object is located; when the loss value is greater than a preset threshold, updating the network weights of the neural network model; when the loss value is less than the preset threshold, stopping updating the network weights of the neural network model.

[0179] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0180] Example 10

[0181] According to an embodiment of the present application, a data processing method is also provided.

[0182] Fig.14 is a flow chart of a sixth data processing method according to an embodiment of the present application. Fig.14 As shown, the method comprises the following steps:

[0183] Step S142, receiving data to be processed.

[0184] The data to be processed in the above steps may be voice data that needs to be recognized, and the voice data may be data collected by a mobile device or a robot such as a smart speaker from the voice of a user; the data to be processed may also be an image that needs to be cut out, and the image may be an image taken by a user through a mobile device or a robot such as a smart speaker, or an image searched by a user on a mobile device, but is not limited thereto. In the embodiments of the present application, an image is used as an example for explanation.

[0185] Step S144, obtaining a neural network model.

[0186] The neural network model in the above steps can adopt an encoder-2decoder model structure shared by a feature encoding model, wherein the encoder model is shared and two decoder models are connected externally, namely decoder-one and decoder-two. Decoder-one is used for forward prediction of the target area, while decoder-two is used to complete accurate data processing.

[0187] Step S146, input the data to be processed into the shared encoding network of the neural network model to obtain the target features of the data to be processed.

[0188] The encoding network in the above steps can be as follows Figure 4 The encoder model shown.

[0189] Step S148, input the target feature into the first decoding network of the neural network model to obtain a processing result.

[0190] The processing result in the above steps may be a recognition result obtained by performing voice recognition on the voice data sent by the user, or may be a cutout result obtained by cutting out the image taken or searched by the user, but is not limited thereto.

[0191] In the above embodiment of the present application, the method also includes: obtaining at least one group of training samples, wherein the training samples include: image data, and label information of the area where the object is located contained in the image data; inputting at least one group of training samples into a shared encoding network to obtain feature data of at least one group of training samples; inputting the feature data into a first decoding network to determine a first result; inputting the feature data and output information of at least one feature layer in the first decoding network into a second decoding network of the neural network model to obtain a second result; and updating the network weights of the neural network model based on the first result and the second result.

[0192] In the above embodiment of the present application, updating the network weights of the neural network model based on the processing results of the area and the label information of the area where the object is located includes: obtaining the loss value of the neural network model based on the processing results of the area and the label information of the area where the object is located; when the loss value is greater than a preset threshold, updating the network weights of the neural network model; when the loss value is less than the preset threshold, stopping updating the network weights of the neural network model.

[0193] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0194] Embodiment 11

[0195] The embodiment of the present application may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal may also be replaced by a terminal device such as a mobile device.

[0196] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of the computer network.

[0197] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the data processing method: receiving an image to be processed, wherein the image to be processed includes a target object; processing the image to be processed using an image segmentation model to determine a target area where the target object is located, wherein the image segmentation model is obtained based on a shared encoding network training.

[0198] Optionally, Fig.15 is a structural block diagram of a computer terminal according to an embodiment of the present application. Fig.15 As shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 1102 and a memory 1104.

[0199] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the data processing method and device in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned data processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0200] The processor can call the information and application stored in the memory through the transmission device to execute the following steps: receiving an image to be processed, wherein the image to be processed includes a target object; using an image segmentation model to process the image to be processed to determine a target area where the target object is located, wherein the image segmentation model is obtained based on a shared encoding network training.

[0201] Optionally, the processor may also execute program codes of the following steps: inputting the image to be processed into a shared encoding network to obtain target features of the image to be processed; and inputting the target features into a first decoding network to obtain a target area.

[0202] Optionally, the processor may also execute the program code of the following steps: obtaining at least one set of training samples, wherein the training samples include: image data, and label information of the area where the object is located contained in the image data; inputting at least one set of training samples into a shared encoding network to obtain feature data of at least one set of training samples; inputting the feature data into a first decoding network to determine the area where the object is located; inputting the feature data and output information of at least one feature layer in the first decoding network into a second decoding network to obtain a processing result of the area; and updating the network weights of the image segmentation model based on the processing result of the area and the label information of the area where the object is located.

[0203] Optionally, the processor may also execute the following program code: obtaining a loss value of the image segmentation model based on the processing result of the region and the label information of the region where the object is located; when the loss value is greater than a preset threshold, updating the network weights of the image segmentation model; when the loss value is less than a preset threshold, stopping updating the network weights of the image segmentation model.

[0204] The processor can call the information and application stored in the memory through the transmission device to execute the following steps: receiving an image to be processed, wherein the image to be processed includes a target object; using an image segmentation model to process the image to be processed to determine a target area where the target object is located, wherein the image segmentation model is obtained based on a shared encoding network training; and displaying an image of the target area.

[0205] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain at least one group of training samples, wherein the training samples include: image data, and label information of the area where the object is located contained in the image data; input at least one group of training samples into the shared encoding network of the image segmentation model to obtain feature data of at least one group of training samples; input the feature data into the first decoding network of the image segmentation model to determine the area where the object is located; input the feature data and the output information of at least one feature layer in the first decoding network into the second decoding network to obtain the processing result of the area; based on the processing result of the area and the label information of the area where the object is located, update the network weights of the image segmentation model.

[0206] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: receive a processing request; based on the processing request, obtain an initial model and at least one set of training samples, wherein the training samples include: image data, and label information of the area where the object is located contained in the image data; based on the encoding network shared in the initial model, use at least one set of training samples to train the initial model to obtain an image segmentation model; output the image segmentation model.

[0207] By using the embodiment of the present application, after acquiring the image to be processed, the image segmentation model can be used to process the image to be processed, and the target area where the target object is located can be determined, thereby achieving the purpose of image cutout on the terminal.

[0208] It is easy to notice that the image segmentation model is trained based on a shared coding network. Compared with the existing technology, the training accuracy of the image segmentation model can be improved through the shared coding network, and in the process of processing the image to be processed, the network time consumption can be saved, thereby achieving the technical effect of saving processing time and improving processing performance, and realizing the needs of quickly completing cutouts or real-time cutouts.

[0209] The solution provided by the embodiment of the present application solves the technical problem in the related art that the processing flow of the image to be processed is highly complex, resulting in a long processing time.

[0210] It can be understood by those skilled in the art that Fig.15 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Fig.15 It does not limit the structure of the above electronic device. For example, the computer terminal A may also include Fig.15 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Fig.15 Different configurations are shown.

[0211] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0212] Example 12

[0213] The embodiment of the present application further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the data processing method provided by the above embodiment.

[0214] Optionally, in this embodiment, the storage medium may be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile devices in a mobile device group.

[0215] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: receiving an image to be processed, wherein the image to be processed includes a target object; processing the image to be processed using an image segmentation model to determine a target area where the target object is located, wherein the image segmentation model is trained based on a shared encoding network.

[0216] Optionally, the storage medium is also configured to store program codes for executing the following steps: inputting the image to be processed into a shared encoding network to obtain target features of the image to be processed; and inputting the target features into a first decoding network to obtain a target area.

[0217] Optionally, the storage medium is also configured to store program codes for executing the following steps: obtaining at least one set of training samples, wherein the training samples include: image data, and label information of the area where the object is located contained in the image data; inputting at least one set of training samples into a shared encoding network to obtain feature data of at least one set of training samples; inputting the feature data into a first decoding network to determine the area where the object is located; inputting the feature data and output information of at least one feature layer in the first decoding network into a second decoding network to obtain a processing result of the area; and updating the network weights of the image segmentation model based on the processing result of the area and the label information of the area where the object is located.

[0218] Optionally, the storage medium is also configured to store program codes for executing the following steps: obtaining a loss value of the image segmentation model based on the processing result of the region and the label information of the region where the object is located; when the loss value is greater than a preset threshold, updating the network weights of the image segmentation model; when the loss value is less than a preset threshold, stopping updating the network weights of the image segmentation model.

[0219] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: receiving an image to be processed, wherein the image to be processed includes a target object; processing the image to be processed using an image segmentation model to determine a target area where the target object is located, wherein the image segmentation model is trained based on a shared encoding network; and displaying an image of the target area.

[0220] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining at least one set of training samples, wherein the training samples include: image data, and label information of the area where the object is located contained in the image data; inputting at least one set of training samples into a shared encoding network of an image segmentation model to obtain feature data of at least one set of training samples; inputting the feature data into a first decoding network of the image segmentation model to determine the area where the object is located; inputting the feature data and output information of at least one feature layer in the first decoding network into a second decoding network to obtain a processing result of the area; and updating the network weights of the image segmentation model based on the processing result of the area and the label information of the area where the object is located.

[0221] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: receiving a processing request; based on the processing request, obtaining an initial model and at least one set of training samples, wherein the training samples include: image data, and label information of an area where an object is located contained in the image data; based on the encoding network shared in the initial model, training the initial model using at least one set of training samples to obtain an image segmentation model; and outputting the image segmentation model.

[0222] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0223] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0224] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0225] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0226] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0227] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.

[0228] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A data processing method, comprising: Receive processing requests; Based on the processing request, an initial model and at least one set of training samples are obtained, wherein the training samples include: image data, and label information of an area where an object is located contained in the image data; Based on the encoding network shared in the initial model, the initial model is trained using the at least one group of training samples to obtain an image segmentation model; Outputting the image segmentation model; The step of training the initial model based on the shared coding network in the initial model using the at least one set of training samples to obtain the image segmentation model comprises: Inputting the at least one set of training samples into the shared encoding network to obtain feature data of the at least one set of training samples; Inputting the feature data into a first decoding network of the initial model to determine the area where the object is located; Inputting the feature data and output information of at least one feature layer in the first decoding network into the second decoding network of the initial model to obtain a processing result of the region; Based on the processing result of the region and the label information of the region where the object is located, the network weights of the initial model are updated to obtain the image segmentation model.

2. The method according to claim 1, wherein: Based on the processing result of the region and the label information of the region where the object is located, updating the network weights of the initial model includes: Obtaining a loss value of the initial model based on a processing result of the region and label information of the region where the object is located; When the loss value is greater than a preset threshold, updating the network weights of the initial model; When the loss value is less than the preset threshold, the updating of the network weights of the initial model is stopped.

3. A data processing method, comprising: Receiving an image to be processed, wherein the image to be processed includes a target object; Processing the image to be processed by using an image segmentation model to determine a target area where the target object is located, wherein the image segmentation model is obtained by training based on a shared encoding network; displaying an image of the target area; Wherein, the method further comprises: Acquire at least one set of training samples, wherein the training samples include: image data, and label information of an area where an object contained in the image data is located; Inputting the at least one set of training samples into the shared encoding network to obtain feature data of the at least one set of training samples; Inputting the feature data into a first decoding network to determine the area where the object is located; Inputting the feature data and output information of at least one feature layer in the first decoding network into a second decoding network to obtain a processing result of the region; Based on the processing result of the region and the label information of the region where the object is located, the network weights of the image segmentation model are updated.

4. The method according to claim 3, wherein: After receiving the image to be processed, the method further includes: Outputting the image segmentation model; Receiving a selected target decoding network, wherein the target decoding network is the first decoding network or the second decoding network in the image segmentation model; Inputting the image to be processed into the shared encoding network to obtain target features of the image to be processed; The target features are processed using the target decoding network to obtain the target area.

5. The method according to claim 4, wherein: Processing the target feature by using the target decoding network to obtain the target area includes: When the target decoding network is the first decoding network, inputting the target feature into the first decoding network to obtain the target region; In the case where the target decoding network is the second decoding network, the target feature is input into the first decoding network, and the target feature and output information of at least one feature layer in the first decoding network are input into the second decoding network to obtain the target area.

6. The method according to claim 3, wherein: Based on the processing result of the region and the label information of the region where the object is located, updating the network weights of the image segmentation model includes: Obtaining a loss value of the image segmentation model based on a processing result of the region and label information of the region where the object is located; When the loss value is greater than a preset threshold, updating the network weights of the image segmentation model; When the loss value is less than the preset threshold, the updating of the network weights of the image segmentation model is stopped.

7. A data processing method, comprising: Receiving an image to be processed, wherein the image to be processed includes a target object; Get the neural network model; Inputting the image to be processed into the shared encoding network of the neural network model to obtain target features of the image to be processed; Inputting the target feature into the first decoding network of the neural network model to obtain the target area where the target object is located; Wherein, the method further comprises: Acquire at least one set of training samples, wherein the training samples include: image data, and label information of an area where an object contained in the image data is located; Inputting the at least one set of training samples into the shared encoding network to obtain feature data of the at least one set of training samples; Inputting the feature data into the first decoding network to determine the area where the object is located; Inputting the feature data and output information of at least one feature layer in the first decoding network into the second decoding network of the neural network model to obtain a processing result of the region; Based on the processing result of the area and the label information of the area where the object is located, the network weights of the neural network model are updated.

8. The method according to claim 7, wherein: Based on the processing result of the region and the label information of the region where the object is located, updating the network weights of the neural network model includes: Obtaining a loss value of the neural network model based on the processing result of the region and the label information of the region where the object is located; When the loss value is greater than a preset threshold, updating the network weights of the neural network model; When the loss value is less than the preset threshold, the updating of the network weights of the neural network model is stopped.

9. A data processing method, comprising: Acquire an image to be processed, wherein the image to be processed includes a target object; Processing the image to be processed by using an image segmentation model to determine a target area where the target object is located, wherein the image segmentation model is obtained by training based on a shared encoding network; Wherein, the method further comprises: Acquire at least one set of training samples, wherein the training samples include: image data, and label information of an area where an object contained in the image data is located; Inputting the at least one set of training samples into the shared encoding network to obtain feature data of the at least one set of training samples; Inputting the feature data into a first decoding network to determine the area where the object is located; Inputting the feature data and output information of at least one feature layer in the first decoding network into a second decoding network to obtain a processing result of the region; Based on the processing result of the region and the label information of the region where the object is located, the network weights of the image segmentation model are updated.

10. The method according to claim 9, wherein: Processing the image to be processed by using an image segmentation model to determine a target area where the target object is located includes: Inputting the image to be processed into the shared encoding network to obtain target features of the image to be processed; The target feature is input into a first decoding network to obtain the target area.

11. The method according to claim 10, wherein: Based on the processing result of the region and the label information of the region where the object is located, updating the network weights of the image segmentation model includes: Obtaining a loss value of the image segmentation model based on a processing result of the region and label information of the region where the object is located; When the loss value is greater than a preset threshold, updating the network weights of the image segmentation model; When the loss value is less than the preset threshold, the updating of the network weights of the image segmentation model is stopped.

12. A data processing method, comprising: Acquire at least one set of training samples, wherein the training samples include: image data, and label information of an area where an object contained in the image data is located; Inputting the at least one set of training samples into a shared encoding network of an image segmentation model to obtain feature data of the at least one set of training samples; Inputting the feature data into a first decoding network of the image segmentation model to determine the area where the object is located; Inputting the feature data and output information of at least one feature layer in the first decoding network into a second decoding network to obtain a processing result of the region; Based on the processing result of the region and the label information of the region where the object is located, the network weights of the image segmentation model are updated.

13. The method according to claim 12, wherein: Based on the processing result of the region and the label information of the region where the object is located, updating the network weights of the image segmentation model includes: Obtaining a loss value of the image segmentation model based on a processing result of the region and label information of the region where the object is located; When the loss value is greater than a preset threshold, updating the network weights of the image segmentation model; When the loss value is less than the preset threshold, the updating of the network weights of the image segmentation model is stopped.

14. The method according to claim 12, wherein: After the network weights of the image segmentation model are updated, the method further includes: Acquire an image to be processed, wherein the image to be processed includes a target object; Inputting the image to be processed into the shared encoding network to obtain target features of the image to be processed; The target feature is input into the first decoding network to obtain the target area where the target object is located.

15. A data processing method, comprising: receiving data to be processed; Get the neural network model; Inputting the data to be processed into the shared encoding network of the neural network model to obtain target features of the data to be processed; Inputting the target feature into the first decoding network of the neural network model to obtain a processing result; Wherein, the method further comprises: Acquire at least one set of training samples, wherein the training samples include: image data, and label information of an area where an object contained in the image data is located; Inputting the at least one set of training samples into the shared encoding network to obtain feature data of the at least one set of training samples; Inputting the feature data into a first decoding network to determine the area where the object is located; Inputting the feature data and output information of at least one feature layer in the first decoding network into a second decoding network to obtain a processing result of the region; Based on the processing result of the area and the label information of the area where the object is located, the network weights of the neural network model are updated.

16. A computer-readable storage medium comprising a stored program, wherein: When the program is running, the device where the computer-readable storage medium is located is controlled to execute the data processing method described in any one of claims 1 to 15.

17. A mobile device comprising a memory and a processor, wherein the processor is used to run a program stored in the memory, wherein: When the program is executed, the data processing method according to any one of claims 1 to 15 is executed.

18. A data processing system comprising: processor; as well as A memory is connected to the processor and is used to provide the processor with instructions for processing the following processing steps: obtaining an image to be processed, wherein the image to be processed includes a target object; processing the image to be processed using an image segmentation model to determine a target area where the target object is located, wherein the image segmentation model is obtained based on training of a shared encoding network; wherein the memory is also used to obtain at least one group of training samples, wherein the training samples include: image data, and label information of an area where the object is located contained in the image data; inputting the at least one group of training samples into the shared encoding network to obtain feature data of the at least one group of training samples; inputting the feature data into a first decoding network to determine an area where the object is located; inputting the feature data and output information of at least one feature layer in the first decoding network into a second decoding network to obtain a processing result of the area; and updating the network weights of the image segmentation model based on the processing result of the area and the label information of the area where the object is located.

Citation Information

Patent Citations

  • Image segmentation method and device and electronic device

    CN109325954A

  • Convolutional neural network road scene classification and road segmentation method

    CN109993082A