Data processing method and device, equipment and storage medium

CN121794701APending Publication Date: 2026-04-03BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-04-03

Smart Images

  • Figure CN121794701A_ABST
    Figure CN121794701A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method and device, equipment and a storage medium. The method includes acquiring a first training video related to a reference robot. For at least one training image in the first training video, at least one data expansion strategy corresponding to the at least one training image is determined, and the at least one training image comprises at least one part of the reference robot and an operation object related to the reference robot. And generating a second training video related to robot training based on the at least one training image and the at least one data expansion strategy. A new training video is generated based on an existing training video by adopting a data expansion strategy corresponding to each training image contained in the training video. In this way, the data size and diversity of the training data are improved, and therefore the operation performance of the target robot is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, equipment and storage medium for data processing TECHNICAL FIELD

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, device, equipment and computer readable storage medium for data processing. BACKGROUND

[0002] In recent years, machine learning models have been rapidly developed and have been widely used in multiple technical fields. For example, a machine learning model can be applied to control a robot. In order to ensure the normal work of the robot, a large amount of training data needs to be used to train the machine learning model used by the robot. However, there are some problems in the current training method of the machine learning model, which affects the performance of the machine learning model.

[0003] SUMMARY

[0004] In a first aspect of the present disclosure, a method for data processing is provided. The method comprises: obtaining a first training video related to a reference robot; determining, for at least one training image in the first training video, at least one data augmentation strategy corresponding to the at least one training image respectively, the at least one training image comprising at least a part of the reference robot and an operation object related to the reference robot; and generating a second training video related to training of the reference robot based on the at least one training image and the at least one data augmentation strategy.

[0005] In a second aspect of the present disclosure, a device for data processing is provided. The device comprises: an obtaining module configured to obtain a first training video related to a reference robot; a determining module configured to determine, for at least one training image in the first training video, at least one data augmentation strategy corresponding to the at least one training image respectively, the at least one training image comprising at least a part of the reference robot and an operation object related to the reference robot; and a generating module configured to generate a second training video related to training of the reference robot based on the at least one training image and the at least one data augmentation strategy.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device comprises at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer readable storage medium is provided. The computer readable storage medium has stored thereon a computer program, which is executable by a processor to implement the method of the first aspect.

[0008] It should be understood that the matters described in this section are not intended to define key or essential features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0009] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the annexed drawings in which:

[0010] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0011] FIG. 2 shows a schematic diagram of one example of a data processing system according to some embodiments of the present disclosure;

[0012] FIG. 3 shows a schematic diagram of one example of a training image according to some embodiments of the present disclosure;

[0013] FIG. 4 shows an architectural diagram of one example of a video generation model according to some embodiments of the present disclosure;

[0014] FIG. 5 shows a flowchart of a process of data processing according to some embodiments of the present disclosure;

[0015] FIG. 6 shows a block diagram of an apparatus for data processing according to some embodiments of the present disclosure; and

[0016] FIG. 7 shows a block diagram of a device capable of implementing embodiments of the present disclosure. DETAILED DESCRIPTION

[0017] It can be understood that, before using the technical solutions disclosed by the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the scene of use, etc. should be informed to the user and the authorization of the user should be obtained in accordance with relevant laws and regulations.

[0018] For example, in response to receiving a user's active request, a prompt message is sent to the user to explicitly prompt the user that the operation requested to be performed will require the acquisition and use of the user's personal information. Thus, the user can voluntarily choose whether to provide personal information to the electronic device, application program, server or storage medium, etc. software or hardware that performs the operation of the technical solutions of the present disclosure.

[0019] As an optional but non-limiting implementation, in response to receiving the active request of the user, the manner of sending the prompt information to the user may be, for example, a pop-up window manner, in which the prompt information may be presented in the form of text. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.

[0020] It can be understood that the above notification and user authorization obtaining process is only illustrative, and does not limit the implementation of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0021] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.

[0022] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, rather, these embodiments are provided to make the present disclosure more thorough and complete. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.

[0023] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. Furthermore, embodiments described in any section / subsection can be combined with any other embodiment described in the same section / subsection and / or a different section / subsection in any manner.

[0024] In this document, unless explicitly stated otherwise, performing a step "in response to A" does not mean that the step is performed immediately after A, but can include one or more intermediate steps.

[0025] In the description of embodiments of the present disclosure, the term "comprising" and similar terms are to be interpreted as open inclusion, i.e. "including but not limited to". The term "based on" is to be interpreted as "at least partially based on". The term "one embodiment" or "the embodiment" is to be interpreted as "at least one embodiment". The term "some embodiments" is to be interpreted as "at least some embodiments". Other explicit and implicit definitions can also be included below. The terms "first", "second", etc. can refer to different or same objects. Other explicit and implicit definitions can also be included below.

[0026] As used herein, the term “model” can learn the association between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes the input and provides the corresponding output by using multiple layers of processing units. In this document, “model” can also be referred to as “machine learning model”, “machine learning network” or “network”, which are used interchangeably herein. One model can further include different types of processing units or networks.

[0027] FIG. 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG. 1, a model 130-1 with pre-training parameter values and a model 130-2 with post-training parameter values can be collectively or individually referred to as a model 130. The model 130 can be implemented in or included in an electronic device 140 and / or an electronic device 150.

[0028] In the environment 100 of FIG. 1, it is desirable to train and use such a machine learning model (i.e., the model 130) that is configured for multiple application environments. For example, in the case where the model 130 is a robot control model, a corresponding action plan can be generated based on the user input robot control instructions and the information of the environment in which the robot is located, to control the robot to perform the control instructions. In the case where the model 130 is an instruction recognition model, a corresponding control instruction can be generated based on the user input instruction and the environmental information of the environment in which the robot is located, to control the robot.

[0029] As shown in FIG. 1, the environment 100 includes an electronic device 140 and an electronic device 150. There can be a model training system in the electronic device 140 and a model application system in the electronic device 150. The upper part of FIG. 1 shows the process of the model training stage, and the lower part shows the process of the model application stage. Before training, the parameter values of the model 130 can have initial values, or can have pre-trained parameter values obtained through a pre-training process. The model 130-1 can be trained via forward propagation and back propagation, and the parameter values of the model 130-1 can be updated and adjusted during the training process. The model 130-2 can be obtained after the training is completed. The parameter values of the obtained model 130-2 have been updated, and based on the updated parameter values, the model 130-2 can be used to implement the song synthesis task in the model application stage.

[0030] During a model training phase, the model 130 can be trained based on a training sample set 110 including a plurality of training samples 112, and using a model training system. Here, each training sample 112 can involve a tuple format. For example, for a robot control task, the training sample 112 can include a training input 120 and a training output 122. The training input 120 in the robot control task can include control instructions and action planning. The training sample 112 including the model input 120 and the model output 122 can be used to train the model 130. Specifically, the training process can be performed iteratively using a large number of training samples. After the training is completed, the model 130 can include knowledge about the task to be processed. During a model application phase, the model 130 (now the model 130 has trained parameter values) can be used to perform the corresponding task. For example, a model input 142 in the robot control task can be received, and a corresponding model output 144 indicating the robot action planning can be output.

[0031] In FIG. 1, the electronic device 140 and the electronic device 150 can be control devices of a reference robot, or servers communicatively connected with the reference robot. In some embodiments, the electronic device 140 and the electronic device 150 can include any computing system having computing capability, such as various computing devices / systems, terminal devices, servers, etc. The terminal device can involve any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. The server includes, but is not limited to, a mainframe, an edge computing node, a computing device in a cloud environment, etc.

[0032] It should be understood that the components and arrangements depicted in the environment 100 of FIG. 1 are only examples and that other components and arrangements can be used to implement the exemplary implementations described in this disclosure. The implementations of this disclosure are not limited in this regard.

[0033] As briefly mentioned above, machine learning models have been rapidly developed and have been widely used in multiple technical fields. In order to ensure the performance of the machine learning model, the machine learning model needs to be trained. However, in some special technical fields, due to the difficulty of obtaining training data, the amount of training data obtained is small and lacks diversity. As a result, it is impossible to obtain sufficient training data to train the machine learning model, which affects the performance of the machine learning model. For example, in the field of robot control, due to the high cost of obtaining real machine data, the amount of data used to train the robot control model is small and lacks diversity. The robot control model can only complete simple tasks in a specific environment, and in the case of complex tasks or complex environments for the robot, the performance of the robot control model will be greatly reduced.

[0034] Embodiments of the present disclosure provide a scheme for data processing. According to various embodiments of the present disclosure, a first training video related to a reference robot is obtained. For at least one training image in the first training video, at least one data augmentation strategy corresponding to the at least one training image is determined, the at least one training image including at least a part of the reference robot and an operation object related to the reference robot. Based on the at least one training image and the at least one data augmentation strategy, a second training video related to robot training is generated.

[0035] By adopting the data augmentation strategy corresponding to each training image contained in the training video, a new training video is generated based on the existing training video. In this way, the amount of training data is improved. Thus, the training effect of the model used by the target robot is improved, thereby improving the operation performance of the robot.

[0036] FIG. 2 shows a schematic diagram of one example of a data processing system 200 according to some embodiments of the present disclosure. In the example of FIG. 2, the data processing system 200 can be implemented in or included in an electronic device.

[0037] The electronic device 140 first obtains a first training video 220 related to a reference robot. The first training video 220 is a video for training the reference robot, and the first training video 220 includes a plurality of training videos. In some embodiments, the first training video 220 can be determined according to video data obtained from a public data set and / or the Internet. The first training video 220 is determined according to the application scenario of the reference robot to further improve the training effect of the model used by the reference robot. For example, if the reference robot is a robot applied to a home environment, the first training video 220 is a video containing an indoor scene. If the reference robot is a robot applied to a cargo transportation scenario, the first training video 220 is a video containing a cargo warehouse scene.

[0038] In some embodiments, the first training video 220 can be determined according to photos acquired from public data sets and / or the Internet. The electronic device 140 first acquires photos related to the reference robot, and takes the acquired photos as initial images. The electronic device 140 generates an initial image set including a plurality of initial images by copying the initial images related to the reference robot. Affine transformation operations are respectively performed on the plurality of initial images in the initial image set. Based on the initial image set after affine transformation, the first training video 220 is generated. In some embodiments, the affine transformation on the initial images in the initial image set can be a translation operation, a scaling operation, or a rotation operation, etc. performed on the initial images.

[0039] In some embodiments, the electronic device 140 can also determine the first training video 220 according to real machine data collected during the operation of the robot. For example, data collected by image acquisition sensors or depth sensors during the execution of a path finding task by the robot and action data during the execution of the task by the robot.

[0040] In some embodiments, in order to improve the training effect of the robot control model 250, the collected robot real machine data 211 and Internet data 212 can be provided to the data preprocessing module 210.

[0041] The data preprocessing module 210 is implemented in or included in the electronic device 140. The data preprocessing module 210 is configured to filter and preprocess the provided robot real machine data 211 and Internet data 212 to generate the first training video 220. In some embodiments, the data preprocessing module 210 is configured to eliminate invalid data in the robot real machine data 211 and Internet data 212. For example, portrait images, close-up photos of operation objects, and videos containing split shots.

[0042] For at least one training image in the first training video 220, at least one data augmentation strategy corresponding to the at least one training image is determined. FIG. 3 shows a schematic diagram of one example of a training image according to some embodiments of the present disclosure. As shown in FIG. 3, the training image 360 includes at least a part of the reference robot 310 and an operation object 320 related to the reference robot 310. In some embodiments, the training image 360 further includes an image background 330.

[0043] For each frame training image 360 in the first training video 220, a data augmentation strategy corresponding to the training image 360 is determined. Specifically, the data augmentation strategy includes background replacement (i.e., replacing all or part of the background of the training image 360 with other content), background object addition (i.e., adding other objects to the existing background of the training image 360), background object deletion or modification (i.e., changing the objects in the existing background of the training image 360), operation object replacement (i.e., changing the operation object 320 in the training image 360), video editing, and the like. In some embodiments, the same data augmentation strategy or different data augmentation strategies can be selected for each frame training image 360 in any training video. Alternatively or additionally, a plurality of data augmentation strategies can correspond to any frame training image 360. For example, for a training image 360 including a background A1 and an operation object B1, only the background A1 of the training image 360 can be replaced with A2, or both the background A1 and the operation object B1 of the training image 360 can be replaced with a background A2 and an operation object B2. In some embodiments, the background and operation object 320 used for replacement can be determined from a predetermined object library (e.g., a container word library or an object word library). The object library includes images and representations of at least one augmented object for augmenting the first training video 220.

[0044] Based on the training task, the electronic device 140 determines the data requirement of the trained robot. Based on the data requirement and the first training video 220, the data augmentation strategy is determined. In the following, the robot trained by the second training video will be referred to as the target robot.

[0045] In some embodiments, the corresponding data augmentation strategy is determined based on the type of machine learning model used by the target robot and the task performed. Specifically, the corresponding data augmentation strategy can be determined according to the task scenario in which the target robot is located. For example, if the task scenario of the target robot is a cargo handling scenario, the data augmentation strategy of operation object 320 replacement can be used. Alternatively or additionally, the first training video 220 can be determined according to the data requirement of the target robot. For example, if the task scenario of the target robot is a home environment, the type of data required by the machine learning model used by the target robot can be determined. According to the predetermined correspondence between the type of data and the data augmentation strategy, the corresponding data augmentation strategy is determined. In some embodiments, if there are a large number of training images 360, the training images 360 can be divided into a plurality of groups of training images 360 in proportion, and different data augmentation strategies can be applied to each group of training images 360.

[0046] Based on the training image 360 and one or more data augmentation strategies corresponding to the training image 360, the electronic device 140 generates a second training video 240 related to the training of the target robot.

[0047] In some embodiments, the second training video 240 can be generated by using the machine learning model. The electronic device 140 provides the first training video 220 to the video generation module 230 to obtain the second training video 240 output by the video generation model 230. Specifically, for any training image 360 in the first training video 220, after determining the data augmentation strategy corresponding to the training image 360, the mask region 340 and the description text 350 corresponding to the training image 360 are determined based on the determined data augmentation strategy. Based on the mask region 340 and the description text 350 corresponding to at least one training image 360, the second training video 240 is generated by using the video generation model 230 based on at least one training image 360. Specifically, the description text 350 is used to indicate the operation to be performed on the mask region 340.

[0048] In some embodiments, the mask region 340 is a region determined according to the data augmentation strategy. Specifically, the process of determining the mask region 340 corresponding to the training image 360 based on the data augmentation strategy is described below. The region information of the region where the operation object 320 in the training image 360 is located is determined. Based on the region information, the variable region in the training image 360 is determined. Based on the data augmentation strategy corresponding to the training image 360, the mask region 340 is determined in the variable region. Specifically, the region where the operation object 320 in the training image 360 is located is determined first. The non-variable region in the training image 360 is obtained by superimposing the regions where each operation object 320 in the training image 360 is located, and the other regions are taken as the variable region.

[0049] In some embodiments, the description text 350 is determined based on the image and the identification of the augmented object in the object library. The identification of the augmented object includes the name, the category or the description of the augmented object.

[0050] In some embodiments, if the data augmentation strategy of “background replacement” is adopted for the training image 360, a random region in the training image 360 can be taken as the mask region 340, and the corresponding description text 350 is “complete the mask region 340”. Alternatively or additionally, a container name is randomly selected from the object library as a supplement to the video content, for example, “complete the mask region 340, and the mask region 340 is a water pool”.

[0051] In some embodiments, if the data augmentation strategy of “background object addition” is adopted for the training image 360, a rectangular mask region 340 of a random size can be generated in a random region in the training image 360, and an operation object B1 is randomly selected from the object library, and the corresponding description text 350 is “add an operation object B1 in the mask region 340”.

[0052] In some embodiments, if the data augmentation strategy of “background object deletion or modification” is adopted for the training image 360, the first frame of the first training video 220 can be detected to obtain one or more operation objects 320 in the first frame training image 360. The mask region 340 is determined based on the position of each detected operation object 320. If it is object deletion, the corresponding description text 350 is “complete the mask region 340”. If it is object modification, the description text 350 is “add operation object B1 in the mask region 340”.

[0053] In some embodiments, if the data augmentation strategy of “operation object replacement” is adopted for the training image 360, the mask region 340 can be determined in the non-variable region in the training image 360. Specifically, the operation object 320 in the non-variable region is determined, and the region where the operation object 320 is located is taken as the mask region 340. For example, if the task performed by the reference robot 310 is “opening the drawer”, the drawer region is determined as the mask region 340. The corresponding description text 350 is “add operation object B1 in the mask region 340”. In order to ensure the consistency of the robot action in different training images 360, the operation object 320 before and after replacement can be the same object with different textures and colors.

[0054] In some embodiments, if the data augmentation strategy of “video editing” is adopted for the training image 360, there is no mask region 340. In the case where a container C in the container library is detected from the first frame of the video, an operation object B1 is randomly extracted from the object library, and the corresponding description text 350 is “add operation object B1 in container C”.

[0055] In some embodiments, the process of determining the operation object 320 in the training image 360 is described below. The action of the reference robot 310 corresponding to the training image 360 is determined. Based on the action and the parameters of the image acquisition device, the object region of the operation object 320 in the training image 360 is determined. Based on the object region, the region information is determined by using the segmentation model. Specifically, according to the action sequence of the reference robot 310 and the parameters (such as image capture angle, etc.) of the image acquisition device used to capture the training image 360, the projection point (i.e. the object region) of the reference robot 310 or part of the reference robot 310 in the training image 360 is determined. The seed point is determined near the projection point, and the seed point is provided to the segmentation model to determine the region information of the object region.

[0056] Alternatively or additionally, a semantic segmentation operation can be performed on the training image 360 to obtain planar elements (e.g., a tabletop, a wall, a sink, a drawer, a cabinet, etc.) with an area greater than an area threshold in the training image 360. A mask region 340 is determined in an area where the planar elements are located to further improve the reliability of the augmented data.

[0057] The video generation model 230 generates a second training video 240 from the provided training image 360 and the mask region 340 corresponding to the training image 360, and the description text 350 to train the target robot. Training the target robot indicates training the robot control model 250 adopted by the target robot. In some embodiments, to ensure the training effect of the robot control model, the trained target robot can be a robot with the same task scenario or similar use as the reference robot 310.

[0058] As an example, the video generation model 230 can be a diffusion model constructed based on a Transformer. The output sequence of the video generation model 230 includes the training image 360, a reference image related to the training image 360, the description text 350 related to the training image 360, and robot action information related to the training image 360. The reference image is the training image 360 after adding the mask region 340.

[0059] In some embodiments, first, a reference image corresponding to the training image 360 is determined based on the training image 360 and the mask region 340 corresponding to the training image 360. A training sample sequence is generated based on the training image 360 and the reference image corresponding to the training image 360, at least one action of the reference robot 310, and at least one description text 350. The training sample sequence is provided to the video generation model 230 to generate the second training video 240.

[0060] Since the training video is directly obtained from public data, there can be a problem of too long video duration. Training the robot control model using a too long video can affect the training effect of the model. In some embodiments, first, the duration of the first training video 220 is determined to determine whether the duration exceeds a duration threshold. If the duration of the first training video 220 exceeds the duration threshold, the first training video 220 is split into multiple sub-training videos. The multiple sub-training videos are expanded to increase the amount and diversity of the training video data. Each sub-training video is a continuous video, and the first frame of any sub-training video is the last frame of the previous sub-training video. In response, the mask region 340 in different sub-training videos is consistent.

[0061] During the augmentation of the first training video, the position of the operating object in different training images can change. This situation will affect the consistency between training images, resulting in a lower video authenticity. Therefore, the video generation model 230 needs to be trained to ensure the consistency of the video generated by the video generation model 230. FIG. 4 shows an architecture diagram of one example of a video generation model training system 400 according to some embodiments of the present disclosure. The video generation model 230 can be implemented or included in the electronic device 140. As shown in FIG. 4, the video generation model training system adjusts the parameters of the video generation model based on a loss function 450.

[0062] In some embodiments, the video generation model 230 can be trained in the following way. A fourth training video is obtained by modifying the third training video 410 according to the data augmentation strategy 420. The fourth training video is processed by the video generation model 230 to obtain a reconstruction result 440 corresponding to the third training video 410. By comparing the reconstruction result 440 with the third training video 410, training feedback information is generated. The parameters of the video generation model 230 are updated based on the training feedback information.

[0063] In some embodiments, based on the third training video 410 provided by the Internet data 212 or the public data set, a model training sequence 430 including the fourth training video is obtained. Specifically, first, the third training video 410 is processed according to the difference of data and labels to obtain the model training sequence 430. The form of the model training sequence 430 is a data pair including the third training video 410 (including at least one model training image 431), the fourth training video (i.e. the modified third training video 410, including a plurality of model reference images 432), the training description text 433, the training action information 434 related to the action of the robot in the model training image 431 and the depth map.

[0064] According to the third training video 410 provided to the video generation model 230, a reconstruction result 440 corresponding to the third training video 410 is generated. Based on the loss function 450, the third training video 410 and the reconstruction result 440 are compared to obtain training feedback information. The parameters provided by the video generation model 230 are updated according to the training feedback information.

[0065] Specifically, the fourth training video and the training description text 433 can be determined according to the third training video 410 in the following way. If the third training video 410 does not have a label or a text description, a training mask area is set in the video. The shape of the training mask area is a random pattern. The training mask area is black, and the video with the training mask area added is the fourth training video. The text description is similar to “complete the training mask area”.

[0066] If the third training video 410 has a label or description (e.g., a description about the object category in the video and the use of the video), a training mask region is set in the video, and the text description is similar to "complete the training mask region, and the video describes a person picking up the operation object B1".

[0067] If the third training video 410 contains object detection box or semantic segmentation annotation, a training mask region is generated around the object detection box or semantic segmentation region, and the text description is "add the operation object B1 in the training mask region".

[0068] If the data itself contains a set of video pairs (e.g., high-quality expanded data after manual screening), the video does not need to add a training mask region, and the text description is "add the operation object B1 in the background A1" and "modify the operation object B1 to red".

[0069] If the third training video 410 has depth information, the model training sequence 430 is generated based on the foregoing scheme in the following manner. For example, the fourth training video can contain RGB information and depth information, and the RGB information and the depth correspond to the same training mask region. The fourth training video can contain only RGB information with a training mask region, and the entire depth image is all the training mask region. The fourth training video can contain only depth information, and the RGB information of the image is all the training mask region. The fourth training video can contain only depth information with a training mask region, and the entire RGB information is all the training mask region.

[0070] In some embodiments, if the third training video 410 is the robot real data 211, the training action information 434 is determined based on the pose transformation of the reference robot 310 in each frame of the third training video 410 compared to the previous frame.

[0071] In the case of training the video generation model 230, the model training sequence 430 at least includes the RGB information and the depth information of each model training image 431 in the third training video 410, the RGB information and the depth information of each model training image 431 after adding noise, the RGB information and the depth information of each model training image 431 in the fourth training video, the training description text 433 of each model training image 431, and the training action information 434 contained in the model training image 431.

[0072] In some embodiments, the data filtering module can be used to filter the second training video 240 generated by the video generation module 230 to select the second training video 240 with higher quality. The selected second training video 240 is then provided to the video generation model 230 to further train the video generation model 230 and improve its performance.

[0073] In some embodiments, the video generation model 230 may use different types of data (pure images, pure videos, images / videos with semantic segmentation annotations, images / videos with text annotations, robot operation data, image editing / video editing data, etc.) as training samples to increase the amount and diversity of training data.

[0074] In this way, the consistency of training images within the same training video is ensured through masked regions and descriptive text, thereby improving training effectiveness. On the other hand, the motion information of the reference robot in the training images is preserved to ensure that the generated data can be used to train the robot control model.

[0075] Figure 5 illustrates a flowchart of a data processing procedure 500 according to some embodiments of the present disclosure. The procedure 500 may be implemented or included at an electronic device 140.

[0076] In box 510, obtain the first training video related to the reference robot.

[0077] In some embodiments, an initial image set comprising multiple initial images is generated by copying initial images associated with a reference robot; affine transformation operations are performed on the multiple initial images in the initial image set respectively; and a first training video is generated based on the affine transformed initial image set.

[0078] In box 520, for at least one training image in the first training video, at least one data augmentation strategy corresponding to each of the at least one training image is determined, wherein the at least one training image includes at least a portion of the reference robot and an operation object related to the reference robot.

[0079] In some embodiments, determining a data augmentation strategy includes: determining data requirements for training a reference robot based on a training task; and determining a data augmentation strategy based on the data requirements and a first training video.

[0080] In some embodiments, the data augmentation strategy includes at least one of the following: changing the background of at least one training image, or changing the object of operation in at least one training image.

[0081] In box 530, a second training video related to robot training is generated based on at least one training image and at least one data augmentation strategy.

[0082] In some embodiments, generating the second training video comprises: for a training image in the at least one training image, determining a mask region and a description text corresponding to the training image based on the data augmentation strategy corresponding to the training image, the description text being used to indicate an operation to be performed on the mask region; and generating the second training video based on the mask region and the description text corresponding to the at least one training image, the at least one training image, and the video generation model.

[0083] In some embodiments, determining the description text corresponding to the training image based on the data augmentation strategy comprises: obtaining an object library used to augment the first training video, the object library comprising images and labels of at least one augmented object; and determining the description text based on the at least one data augmentation strategy corresponding to the training image and the images and labels of the at least one augmented object.

[0084] In some embodiments, determining the mask region corresponding to the training image based on the data augmentation strategy comprises: determining region information of a region in which an operation object in the training image is located; determining a variable region in the training image based on the region information; and determining the mask region in the variable region based on the data augmentation strategy corresponding to the training image.

[0085] In some embodiments, determining the region information of the region in which the operation object in the training image is located comprises: determining an action of a reference robot corresponding to the training image; determining an object region of the operation object in the training image based on the action and parameters of an image acquisition device; and determining the region information based on the object region and a segmentation model.

[0086] In some embodiments, generating the second training video based on the video generation model comprises: obtaining at least one reference image corresponding to the at least one training image, the at least one reference image being determined based on the training image and the at least one mask region corresponding to the training image; generating a training sample sequence based on the at least one reference image corresponding to the at least one training image respectively, at least one action of a reference robot, at least one description text, and the at least one training image; and inputting the training sample sequence into the video generation model to generate the second training video.

[0087] In some embodiments, the video generation model is trained by: obtaining a fourth training video obtained by modifying the third training video according to the data augmentation strategy; processing the fourth training video by the video generation model to obtain a reconstruction result corresponding to the third training video; generating training feedback information by comparing the reconstruction result and the third training video; and updating parameters of the video generation model based on the training feedback information.

[0088] In some embodiments, the process 500 further includes, in response to the length of the first training video exceeding a length threshold, splitting the first training video into a plurality of sub-training videos; and generating a second training video based on the plurality of sub-training videos.

[0089] In some embodiments, the process 500 further includes training, based on the training video, a target model to be used by the reference robot.

[0090] FIG. 6 shows a schematic structural block diagram of an apparatus 600 for data processing, according to certain embodiments of the present disclosure. The apparatus 600 can be implemented as or included in the electronic device 140. Various modules / components in the apparatus 600 can be implemented by hardware, software, firmware, or any combination thereof.

[0091] As shown, the apparatus 600 includes an obtaining module 610 configured to obtain a first training video related to a reference robot. The apparatus 600 further includes a determining module 620 configured to determine, for at least one training image in the first training video, at least one data augmentation strategy corresponding to the at least one training image respectively, the at least one training image including at least one part of the reference robot and an operation object related to the reference robot. The apparatus 600 further includes a generating module 630 configured to generate a second training video related to robot training based on the at least one training image and the at least one data augmentation strategy.

[0092] In some embodiments, the obtaining module 610 is further configured to generate an initial image set including a plurality of initial images by copying an initial image related to the reference robot; perform an affine transformation operation on the plurality of initial images in the initial image set respectively; and generate the first training video based on the initial image set after the affine transformation.

[0093] In some embodiments, the determining module 620 is further configured to determine a data requirement for training the reference robot based on a training task; and determine the data augmentation strategy based on the data requirement and the first training video.

[0094] In some embodiments, the generating module 630 is further configured to, for a training image in the at least one training image, determine a mask region corresponding to the training image and a description text based on the data augmentation strategy corresponding to the training image, the description text being used to indicate an operation to be performed on the mask region; and generate the second training video based on the mask region and the description text corresponding to the at least one training image, the at least one training image, and a video generation model.

[0095] In some embodiments, the generating module 630 is further configured to obtain an object library for augmenting the first training video, the object library comprising images and labels of at least one augmented object; and determine the description text based on at least one data augmentation strategy corresponding to the training image and the images and labels of the at least one augmented object.

[0096] In some embodiments, the generating module 630 is further configured to determine region information of a region in which the operation object is located in the training image; determine a variable region in the training image based on the region information; and determine a mask region in the variable region based on the data augmentation strategy corresponding to the training image.

[0097] In some embodiments, the generating module 630 is further configured to determine an action of a reference robot corresponding to the training image; determine an object region of the operation object in the training image based on the action and parameters of the image acquisition device; and determine the region information using the segmentation model based on the object region.

[0098] In some embodiments, the generating module 630 is further configured to obtain at least one reference image corresponding to the at least one training image, the at least one reference image being determined based on the training image and at least one mask region corresponding to the training image; generate a training sample sequence based on the at least one reference image corresponding to the at least one training image respectively, at least one action of the reference robot, at least one description text, and the at least one training image; and input the training sample sequence into the video generation model to generate the second training video.

[0099] In some embodiments, the generating module 630 is further configured to obtain a fourth training video obtained by modifying the third training video according to the data augmentation strategy; process the fourth training video using the video generation model to obtain a reconstruction result corresponding to the third training video; generate training feedback information by comparing the reconstruction result and the third training video; and update parameters of the video generation model based on the training feedback information.

[0100] In some embodiments, the apparatus 600 further comprises a video splitting module configured to split the first training video into a plurality of sub-training videos in response to a time length of the first training video exceeding a time length threshold; and generate the second training video based on the plurality of sub-training videos.

[0101] In some embodiments, the apparatus 600 further comprises a robot model training module configured to train a target model to be used by the reference robot based on the training video.

[0102] FIG. 7 illustrates a block diagram of an electronic device 700 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 700 illustrated in FIG. 7 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 700 illustrated in FIG. 7 can be used to implement the electronic device 140 of FIG. 1.

[0103] As illustrated in FIG. 7, the electronic device 700 is in the form of a general electronic device. Components of the electronic device 700 can include, but are not limited to, one or more processors or processing units 710, a memory 720, a storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 can be a real or virtual processor and capable of executing various processing according to programs stored in the memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve parallel processing capability of the electronic device 700.

[0104] The electronic device 700 typically includes a number of computer storage media. Such media can be any available media that is accessible by the electronic device 700 and includes both volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 can be a removable or non-removable media and can include a machine-readable medium, such as a flash drive, a magnetic disk drive, or any other medium that can be used to store information and / or data and that can be accessed by the electronic device 700.

[0105] The electronic device 700 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7, a disk drive and a disk drive interface can be provided for reading from or writing to a removable, non- volatile magnetic disk (e.g., a "hard disk that can be used for storing software and / or data) and an optical disk drive and an optical disk drive interface can be provided for reading from or writing to a removable, non-volatile optical disk (such as a CD-ROM or other optical medium). In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 720 can include a computer program product 725 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.

[0106] The communication unit 740 enables communication through communication media with other electronic devices. Additionally, the functionality of the components of the electronic device 700 can be implemented in a single computing cluster or a plurality of computer machines capable of communicating over a communication connection. As such, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0107] The input device 750 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 760 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 700 can also communicate with one or more external devices (not shown), such as a storage device, a display device, etc., through the communication unit 740, as needed, with one or more devices that enable a user to interact with the electronic device 700, or with any devices (e.g., a network card, a modem, etc.) that enable the electronic device 700 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0108] According to an example implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, where the computer-executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, there is also provided a computer program product tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, where the computer-executable instructions are executed by a processor to implement the method described above.

[0109] Various aspects of the disclosure can be described in the context of flow diagrams and / or block diagrams that illustrate the functions and / or acts performed by the methods, apparatus, devices, and computer program products according to this disclosure. It is to be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by computer readable program instructions.

[0110] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flow diagrams and / or block diagrams. These computer readable program instructions can also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flow diagrams and / or block diagrams.

[0111] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0112] The flow diagrams and the block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various implementations of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions (s). In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in some cases, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and

[0113] implementations. Numerous modifications and adaptations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The herein disclosed subject matter is to be considered merely illustrative in nature and is not intended to limit the scope of the described implementations as set forth in the appended claims. Rather, the scope of the described implementations is to be understood only as set forth in the appended claims.

Claims

1. A data processing method, comprising: obtaining a first training video related to a reference robot; determining, for at least one training image in the first training video, at least one data augmentation strategy corresponding to the at least one training image respectively, the at least one training image comprising at least a part of the reference robot and an operation object related to the reference robot; and generating a second training video related to robot training based on the at least one training image and the at least one data augmentation strategy.

2. The method of claim 1, wherein the generating the second training video comprises: determining, for a training image in the at least one training image, a mask region corresponding to the training image and a description text based on the data augmentation strategy corresponding to the training image, the description text being used to indicate an operation to be performed on the mask region; and generating the second training video based on the mask region and the description text corresponding to the at least one training image, and the at least one training image, using a video generation model.

3. The method of claim 2, wherein the determining the description text corresponding to the training image based on the data augmentation strategy comprises: obtaining an object library used to augment the first training video, the object library comprising images and identities of at least one augmented object; and determining the description text based on the at least one data augmentation strategy corresponding to the training image and the images and the identities of the at least one augmented object.

4. The method of claim 2, wherein the determining the mask region corresponding to the training image based on the data augmentation strategy comprises: determining region information of a region where the operation object in the training image is located; determining a variable region in the training image based on the region information; and determining the mask region in the variable region based on the data augmentation strategy corresponding to the training image.

5. The method of claim 4, wherein the determining the region information of the region where the operation object in the training image is located comprises: determining a motion of the reference robot corresponding to the training image; determining an object region of the operation object in the training image based on the motion and parameters of an image acquisition device; and determining the region information using a segmentation model based on the object region.

6. The method of claim 1, further comprising: splitting the first training video into a plurality of sub-training videos in response to a time length of the first training video exceeding a time length threshold; and generating the second training video based on the plurality of sub-training videos.

7. The method of claim 2, wherein the generating the second training video using the video generation model comprises: obtaining at least one reference image corresponding to the at least one training image, the at least one reference image being determined based on the training image and the at least one mask region corresponding to the training image; generating a training sample sequence based on the at least one reference image corresponding to the at least one training image respectively, at least one motion of the reference robot, at least one description text, and the at least one training image; and inputting the training sample sequence into the video generation model to generate the second training video.

8. The method of claim 1, wherein the obtaining the first training video comprises: generating an initial image set comprising a plurality of initial images by copying an initial image related to the reference robot; and generating the first training video based on the initial image set. ​ ​ ​ ​ ​ ​ performing an affine transformation operation on each of the plurality of initial images in the initial image set; and generating a first training video based on the set of affine-transformed initial images.

9. The method of claim 2, wherein the video generation model is trained by: obtaining a fourth training video by modifying the third training video according to a data augmentation strategy; processing the fourth training video using the video generation model to obtain a reconstruction result corresponding to the third training video; generating training feedback information by comparing the reconstruction result and the third training video; and updating parameters of the video generation model based on the training feedback information.

10. The method of claim 1, wherein determining the data augmentation strategy comprises: determining a data requirement for training the reference robot based on the training task; and determining the data augmentation strategy based on the data requirement and the first training video.

11. The method of claim 10, wherein the data augmentation strategy comprises at least one of: changing a background of at least one training image, or changing an operation object in at least one training image.

12. The method of claim 1, further comprising: training a target model to be used by the reference robot based on the training video.

13. An apparatus for data processing, comprising: an obtaining module configured to obtain a first training video related to a reference robot; a determining module configured to determine, for at least one training image in the first training video, at least one data augmentation strategy corresponding to the at least one training image, the at least one training image including at least a portion of the reference robot and an operation object related to the reference robot; and a generating module configured to generate a second training video related to robot training based on the at least one training image and the at least one data augmentation strategy.

14. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1-12.

15. A computer-readable storage medium having stored thereon a computer program, the computer program executable by a processor to implement the method according to any one of claims 1-12. ​ ​ ​ ​