Processing method, terminal device, and storage medium
Patent Information
- Application Number
- CN202610968833.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
AI Technical Summary
[0015]本申请通过上述技术方案,能够对图像中的目标对象进行特征提取得到特征提取结果后,将特征提取结果作为衍生物预测模型的条件输入,使得衍生物预测模型执行基于条件驱动的衍生物的推理预测以实现智能、精确地定位目标对象的衍生物,从而获取到包括目标对象及其衍生物的待处理对象后能够对待处理对象中的目标对象及其衍生物进行协同处理,以保障处理结果达到自然、完整的处理效果。因此,本申请的技术方案能够自动、准确地预测并确定目标对象所关联的衍生物,以实现对目标对象及其衍生物的协同处理,保障处理结果达到自然、完整的处理效果,从而提升用户体验。
Smart Images

Figure CN122820899A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, specifically to a processing method, terminal device, and storage medium. Background Technology
[0002] In image processing, it is often necessary to identify, locate, and subsequently edit specific target objects in an image. In real-world scenarios, target objects are often accompanied by various derivatives, such as: afterimages left by moving objects, fluid diffusions such as smoke or flames, debris or splashes generated during object interactions, and optical phenomena such as shadows, reflections, and reflections of the target object.
[0003] When editing a target object in an image is required, the relevant technology has difficulty automatically identifying and locating derivatives associated with the target object. Therefore, it only processes the target object itself and ignores its derivatives, resulting in unnatural processing results and obvious residual traces. Summary of the Invention
[0004] The technical problem this application aims to solve is how to automatically and accurately predict and determine the derivatives associated with a target object based on its characteristics, so as to achieve collaborative processing of the target object and its derivatives. To address the above technical problem, this application provides a processing method, a terminal device, and a storage medium.
[0005] In a first aspect, this application provides a processing method, comprising: performing derivative prediction of the target object based on the feature extraction result of the target object in the image information and the derivative prediction model to obtain the object to be processed in the image information, the object to be processed including the target object and derivatives; and performing image processing on the object to be processed in the image information.
[0006] In one embodiment, image processing includes at least one of the following: Editing or processing a portion or the entirety of the object to be processed in a frame of a video; Editing or processing a part or the whole of the object to be processed in the image; Editing processes include at least one of the following: removal, blurring, and image quality optimization.
[0007] In one implementation, the method for determining the target object includes one of the following: In response to a selection operation on at least a portion of a frame of a video, the selected area is taken as the target object; In response to a selection operation on at least a portion of an image, the selected area is used as the target object.
[0008] In one implementation, the feature extraction strategy includes at least one of the following: Feature extraction is performed on the target object to obtain its feature information; Feature extraction is performed on images containing target objects to obtain feature information of the image. The feature information of the target object and the feature information of the image are fused to obtain the fused feature information. The feature extraction result is determined based on the feature information of the target object or the fused feature information; Feature fusion can take at least one of the following forms: cross-attention, feature multiplication, and feature concatenation.
[0009] In one implementation, feature extraction of the target object includes at least one of the following: Perform object segmentation based on the target object and the image segmentation model, and obtain the image mask of the target object; Spatial and / or semantic features are extracted from the image mask of the target object to obtain the feature information of the target object; Spatial and / or semantic features are extracted from the target object in the image information to obtain the feature information of the target object.
[0010] In one implementation, the derivative prediction model is constructed in a manner that includes at least one of the following: Construct multiple training sample data; The base model is trained based on the constructed training sample data to build a derivative prediction model; The basic model is built from convolutional networks and / or Transformer networks; The loss function used in model training is obtained by combining the binary cross-entropy loss function and the boundary optimization loss function. The learning rate scheduling mechanism used in model training includes cosine annealing or linear descent.
[0011] In one embodiment, the processing method further includes: displaying image information; highlighting the object to be processed in the displayed image information; and adjusting and updating the object to be processed in response to receiving an editing operation on the object to be processed.
[0012] In one embodiment, the processing method further includes at least one of the following: Based on the objects to be processed before the update, obtain negative training sample data; Based on the updated object to be processed, obtain the positive training sample data; The derivative prediction model is adjusted by processing the model parameters based on negative training sample data and / or positive training sample data.
[0013] Secondly, this application also provides a terminal device, including a memory and a processor, wherein the memory stores a processing program or instructions, and when the processing program or instructions are executed by the processor, they implement the steps of the method described above.
[0014] Thirdly, this application also provides a storage medium storing a computer program or instructions, which, when executed by a terminal device, implement the steps of the processing method described above.
[0015] This application, through the aforementioned technical solution, enables feature extraction of target objects in an image. The extracted features are then used as input to a derivative prediction model, allowing the model to perform condition-driven derivative inference and prediction. This achieves intelligent and precise localization of derivatives associated with the target object. By acquiring a process object that includes both the target object and its derivatives, the application can collaboratively process these elements, ensuring a natural and complete processing result. Therefore, this application's technical solution can automatically and accurately predict and determine the derivatives associated with a target object, enabling collaborative processing of the target object and its derivatives, ensuring a natural and complete processing result, and thus improving the user experience. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without any creative effort.
[0017] Figure 1 This is a schematic diagram of the hardware structure of a mobile terminal provided in an embodiment of this application.
[0018] Figure 2 This is a communication network system architecture diagram provided for an embodiment of this application.
[0019] Figure 3 This is a flowchart illustrating a processing method provided in an embodiment of this application.
[0020] Figure 4 This is a flowchart illustrating another processing method provided in an embodiment of this application.
[0021] The realization of the objectives, functional features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0022] It should be understood that although the terms "first," "second," "third," etc., may be used herein to describe various information, these terms are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. The terms "or," "and / or," and "including at least one of the following," as used in this application, can be interpreted as inclusive or meaning any one or any combination thereof. For example, "including at least one of the following: A, B, C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C," and similarly, "A, B, or C" or "A, B, and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C." Exceptions to this definition only occur when combinations of elements, functions, steps, or operations are inherently mutually exclusive in some manner.
[0023] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0024] It should be noted that step designations such as S11 and S12 are used in this document for the purpose of more clearly and concisely describing the corresponding content, and do not constitute a substantial limitation on the order. In specific implementation, those skilled in the art may execute S12 first and then S11, etc., but these should all be within the protection scope of this application.
[0025] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0026] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0027] Terminal devices can be implemented in various forms. For example, the terminal devices described in this application may include terminal devices such as mobile phones, tablets, laptops, handheld computers, personal digital assistants (PDAs), portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, and fixed terminals such as digital TVs and desktop computers.
[0028] The following description will use a mobile terminal as an example. Those skilled in the art will understand that, apart from elements specifically designed for mobile purposes, the construction according to the embodiments of this application can also be applied to fixed-type terminals.
[0029] Please see Figure 1 This is a schematic diagram of the hardware structure of a mobile terminal implementing various embodiments of this application. The mobile terminal 100 may include: an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (Audio / Video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111, etc. Those skilled in the art will understand that... Figure 1 The mobile terminal structure shown does not constitute a limitation on the mobile terminal. The mobile terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0030] The following is combined Figure 1 A detailed introduction to each component of the mobile terminal: The radio frequency unit 101 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with the processor 110; additionally, it transmits uplink data to the base station. Typically, the radio frequency unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, and a duplexer. Furthermore, the radio frequency unit 101 can also communicate wirelessly with networks and other devices. The aforementioned wireless communications may use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), TDD-LTE (Time Division Duplexing-Long Term Evolution), 5G, and 6G.
[0031] WiFi is a short-range wireless transmission technology. Mobile terminals, through the WiFi module 102, can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 1 WiFi module 102 is shown, but it is understood that it is not a necessary component of a mobile terminal and can be omitted as needed without changing the nature of the invention.
[0032] The audio output unit 103 can convert audio data received by the radio frequency unit 101 or the WiFi module 102 or stored in the memory 109 into audio signals and output them as sound when the mobile terminal 100 is in call signal receiving mode, call mode, recording mode, voice recognition mode, broadcast receiving mode, etc. Furthermore, the audio output unit 103 can also provide audio output related to specific functions performed by the mobile terminal 100 (e.g., call signal receiving sound, message receiving sound, etc.). The audio output unit 103 may include a speaker, a buzzer, etc.
[0033] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames can be displayed on the display unit 106. The image frames processed by the GPU 1041 can be stored in the memory 109 (or other storage media) or transmitted via the radio frequency unit 101 or the WiFi module 102. The microphone 1042 can receive sound (audio data) in operating modes such as telephone call mode, recording mode, and voice recognition mode, and can process such sound into audio data. The processed audio (voice) data can be converted into a format that can be transmitted to a mobile communication base station via the radio frequency unit 101 in telephone call mode. The microphone 1042 can implement various types of noise cancellation (or suppression) algorithms to eliminate (or suppress) noise or interference generated during the reception and transmission of audio signals.
[0034] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Optionally, the light sensor includes an ambient light sensor and a proximity sensor. Optionally, the ambient light sensor can adjust the brightness of the display panel 1061 according to the ambient light level, and the proximity sensor can turn off the display panel 1061 and / or backlight when the mobile terminal 100 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. Other sensors that may be configured in the phone, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0035] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0036] User input unit 107 can be used to receive input numerical or character information, and generate key signal inputs related to user settings and function control of the mobile terminal. Optionally, user input unit 107 may include touch panel 1071 and other input devices 1072. Touch panel 1071, also known as touch screen, can collect touch operations on or near the user (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near touch panel 1071), and drive corresponding connection devices according to a pre-set program. Touch panel 1071 may include two parts: a touch detection device and a touch controller. Optionally, the touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to processor 110, and can receive and execute commands sent by processor 110. In addition, touch panel 1071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may also include other input devices 1072. Optionally, other input devices 1072 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc., without being specifically limited here.
[0037] Optionally, the touch panel 1071 may cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of touch event. Subsequently, the processor 110 provides corresponding visual output on the display panel 1061 based on the type of touch event. Although in Figure 1 In this embodiment, the touch panel 1071 and the display panel 1061 are two independent components to realize the input and output functions of the mobile terminal. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the mobile terminal. The specific implementation is not limited here.
[0038] Interface unit 108 serves as an interface through which at least one external device can connect to mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, and so on. Interface unit 108 may be used to receive input (e.g., data, power, etc.) from the external device and transmit the received input to one or more elements within mobile terminal 100, or it may be used to transmit data between mobile terminal 100 and the external device.
[0039] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a program storage area and a data storage area. Optionally, the program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory 109 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0040] The processor 110 is the control center of the mobile terminal. It connects various parts of the mobile terminal via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 109, and by calling data stored in the memory 109, it performs various functions and processes data of the mobile terminal, thereby providing overall monitoring of the mobile terminal. The processor 110 may include one or more processing units; preferably, the processor 110 may integrate an application processor and a modem processor. Optionally, the application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 110.
[0041] The mobile terminal 100 may also include a power supply 111 (such as a battery) that supplies power to various components. Preferably, the power supply 111 can be logically connected to the processor 110 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0042] although Figure 1 As not shown, the mobile terminal 100 may also include a Bluetooth module, etc., which will not be described in detail here.
[0043] To facilitate understanding of the embodiments of this application, the communication network system on which the mobile terminal of this application is based is described below.
[0044] Please see Figure 2 , Figure 2 This application provides a communication network system architecture diagram. The communication network system is an LTE system based on the universal mobile communication technology. The LTE system includes a UE (User Equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network) 202, an EPC (Evolved Packet Core) 203, and the operator's IP services 204, which are connected in sequence.
[0045] Optionally, UE201 can be the aforementioned terminal 100, which will not be described in detail here.
[0046] E-UTRAN202 includes eNodeB2021 and other eNodeB2022s. Optionally, eNodeB2021 can connect to other eNodeB2022s via backhaul (e.g., X2 interface). eNodeB2021 connects to EPC203 and can provide UE201 with access to EPC203.
[0047] EPC203 may include an MME (Mobility Management Entity) 2031, an HSS (Home Subscriber Server) 2032, other MMEs 2033, an SGW (Serving Gateway) 2034, a PGW (Packet Data Network Gateway) 2035, and a PCRF (Policy and Charging Rules Function) 2036, etc. Optionally, MME2031 is the control node that handles signaling between UE201 and EPC203, providing bearer and connection management. HSS2032 is used to provide registers to manage functions such as the Home Location Register (not shown in the figure) and stores user-specific information such as service characteristics and data rates. All user data can be sent through SGW2034. PGW2035 can provide UE 201 IP address allocation and other functions. PCRF2036 is the policy and charging control decision point for service data flow and IP bearer resources. It selects and provides available policy and charging control decisions for the policy and charging enforcement function unit (not shown in the figure).
[0048] IP services 204 may include the Internet, intranet, IMS (IP Multimedia Subsystem), or other IP services.
[0049] Although the above description uses the LTE system as an example, those skilled in the art should know that this application is not only applicable to the LTE system, but also to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA, 5G and future new network systems (such as 6G), etc., without limitation.
[0050] Based on the above-described mobile terminal hardware structure and communication network system, various embodiments of this application are proposed.
[0051] First Embodiment Reference Figure 3 , Figure 3 An exemplary flowchart of a processing method is shown. The processing method of this embodiment can be applied to terminal devices (e.g., mobile phones, computers, tablets, etc.), including: S11: Based on the feature extraction results of the target object in the image information and the derivative prediction model, perform derivative prediction of the target object to obtain the object to be processed in the image information. The object to be processed includes the target object and its derivatives. S12: Perform image processing on the object to be processed in the image information.
[0052] In one implementation, the image information includes at least one of video and image.
[0053] Understandably, the target object can refer to a specific object or a specific area of a frame in a video or a picture.
[0054] Understandably, the determination of the target object can be a specific method of automatically or manually specifying or identifying a particular object or a specific image region to be used for derivative prediction and / or subsequent image processing from image information. For example, manually triggering a selection operation on at least a portion of the image to select a specific object or a specific image region as the target object; or, using a specific algorithm or automatic rules to perform selection processing on at least a portion of the image to select a specific object or a specific image region as the target object.
[0055] Understandably, feature extraction results can refer to the feature representation of a target object obtained by extracting features from the target object in image information, used for derivative prediction. This feature representation includes the target object's own attribute information and may selectively incorporate or exclude contextual information from the scene to which the target object belongs. For example, the feature extraction result can be a fused feature representation obtained by extracting and fusing features from both the target object itself and its surrounding scene; or, for instance, the feature extraction result can be a feature representation obtained by extracting features only from the target object itself.
[0056] Understandably, derivative prediction models can be used to take feature extraction results as conditional inputs and perform condition-driven derivative inference predictions to infer and predict derivatives associated with the target object.
[0057] In one implementation, after the derivative prediction model identifies the derivatives associated with the target object, it can further determine their associated attribute information to obtain the derivative prediction result. The attribute information of the derivative includes at least one of the derivative's regional location and morphological attributes.
[0058] In one implementation, the derivative prediction model includes at least one of rule-based reasoning prediction model and end-to-end prediction model based on deep neural networks.
[0059] Among them, the core mechanism of the rule-based reasoning prediction model is the collaboration between preset rules and the logical reasoning engine, which has the advantage of not requiring training data and is easy to implement.
[0060] Among them, the core mechanism of the end-to-end prediction model based on deep neural networks is that the neural network automatically learns a deep mapping between the object and its derivatives.
[0061] In one implementation, the network structure of the end-to-end prediction model based on deep neural networks includes at least one of the following: convolutional networks, Transformer networks, or hybrid network structures obtained by combining multiple neural networks (such as hybrid network structures obtained by combining convolutional networks and Transformer networks).
[0062] In essence, performing image processing on an object in an image means, based on the region positioning information of the object, performing editing processing on the pixel region covered by the object in the image information to change the visual presentation state of the object.
[0063] In one implementation, when the image information is video, editing processing is performed on the pixel region covered by the object to be processed in the image information. For example, editing processing is performed on the pixel region covered by the object to be processed in one frame of the video that includes the target object; or, for example, editing processing is performed on the pixel region covered by the object to be processed or the target object in all frames of the video that include the target object.
[0064] In one embodiment, the derivative includes at least one of the following: shadow, reflection, trail afterimage, smoke, and debris associated with the target object. The shadow is the projection of the target object under illumination; the reflection is the image of the target object on a reflective surface; the trail afterimage is the trajectory left by a moving object; the smoke is a fluid diffuser; and the debris is splashes or fragments generated during object interaction.
[0065] Intuitively, editing can refer to modifying, replacing, or enhancing the pixel content of an object in an image based on the location of that object, in order to change the object's state of existence or perceptual attributes in visual presentation.
[0066] In one implementation, the editing process includes, but is not limited to, at least one of elimination, blurring, and image quality optimization. Elimination can characterize pixel content replacement, changing the object to be processed from visible to invisible, such as background filling or texture compositing. Bluring can characterize pixel value space smoothing / reduction of sharpness to change sharpness into blur, such as Gaussian blur or defocusing. Image optimization can characterize pixel value enhancement to change from basic ordinary image quality to prominent and sharp image quality, such as sharpening adjustment, contrast adjustment, and color adjustment.
[0067] This embodiment, through the above technical solution, can extract features from target objects in an image. After obtaining the feature extraction results, these results are used as conditional inputs to a derivative prediction model. This enables the derivative prediction model to perform condition-driven derivative inference and prediction, achieving intelligent and accurate localization of derivatives of the target object. By acquiring a process object including the target object and its derivatives, the model can collaboratively process the target object and its derivatives within the process object, ensuring a natural and complete processing effect. Therefore, the technical solution of this embodiment can automatically and accurately predict and determine the derivatives associated with a target object, achieving collaborative processing of the target object and its derivatives, ensuring a natural and complete processing effect, and thus improving the user experience.
[0068] Second Embodiment Based on the technical concept of the first embodiment of this application, see [reference] Figure 4 This embodiment provides another processing method, including: S21: In response to a selection operation on at least a portion of the image information, determine the target object; S22: Perform feature extraction processing based on the target object and extraction strategy to obtain the feature extraction results of the target object; S23: Based on the feature extraction results and derivative prediction model, perform derivative prediction of the target object to obtain the object to be processed in the image information. The object to be processed includes the target object and its derivatives. S24: Perform image processing on the object to be processed in the image information.
[0069] The technical solution of this embodiment can determine the user's intention by receiving the user's selection operation in the image information, thereby anchoring the target object (i.e., the core processing subject). Then, through a feature extraction strategy, the target object can be transformed into a feature representation that the machine (derivative prediction model) can accurately understand. After this feature representation is input into the derivative prediction model, the derivative prediction model can expand from the target object to its derivatives to obtain the object to be processed, and thus perform pixel-level processing on the object to be processed. Introducing an extraction strategy after the target object is determined makes feature extraction targeted and directional, so that the extracted feature representation can more accurately and comprehensively reflect the characteristics or attributes of the target object, thereby improving the accuracy of the subsequent derivative prediction model in determining the derivatives of the target object. Furthermore, after determining the object to be processed, which includes the target object and its derivatives, the scope of image processing can be expanded from a single target object to the entirety formed by the target object and its derivatives, thus avoiding processing residues or inconsistencies caused by some implementations that only perform image processing on the target object while ignoring its derivatives.
[0070] In one implementation, in response to a selection operation on at least a portion of a frame in image information, determining a target object (or a method for determining the target object) includes one of the following: In response to a selection operation on at least a portion of a frame of a video, the selected area is taken as the target object; In response to a selection operation on at least a portion of an image, the selected area is used as the target object.
[0071] In one implementation, the selection operation can be a touch operation, voice command, or the like on the user interface. The touch operation includes at least one of the following: clicking, selecting, or swiping on the user interface.
[0072] Thus, the technical solution of this embodiment can equally support selection operations on video frames (temporal media) and images (static media) and perform derivative prediction and image processing, enabling users to enjoy a consistent image processing experience across different types of image information. Through selection operations, a direct, accurate, and flexible mapping channel can be established from the user's visual intent to a specific object. Through close cooperation with downstream derivative prediction, it can achieve an intelligent image processing experience where the user selects a specific object, the system automatically completes the derivative of that object, and performs editing processing.
[0073] Understandingly, an extraction strategy can refer to a set of rules configured for derivative prediction tasks, which regulate the scope of information sources and the combination of features for feature extraction. The scope of information sources includes the target object itself and its associated scene context. The feature combination method includes independent use and / or fusion use to generate feature extraction results adapted to different derivative prediction scenarios.
[0074] In one implementation, the extraction strategy for the feature extraction results includes, but is not limited to, at least one of the following: Feature extraction is performed on the target object to obtain its feature information; Feature extraction is performed on images containing target objects to obtain feature information of the image. The feature information of the target object and the feature information of the image are fused to obtain the fused feature information. The feature extraction result is determined based on the feature information of the target object or the fused feature information.
[0075] Understandably, feature extraction of a target object can refer to the process of extracting and representing spatial and / or semantic features of the region where the target object is located in image information in order to obtain feature information of the target object for derivative prediction.
[0076] In one implementation, feature extraction of the target object includes, but is not limited to, at least one of the following: Perform object segmentation based on the target object and the image segmentation model, and obtain the image mask of the target object; Spatial and / or semantic features are extracted from the image mask of the target object to obtain the feature information of the target object; Spatial and / or semantic features are extracted from the target object in the image information to obtain the feature information of the target object.
[0077] In one implementation, the image segmentation model can be a semantic segmentation model, an instance segmentation model, or a panoptic segmentation model, such as Mask R-CNN, U-Net, DeepLab, etc.
[0078] In one implementation, the image mask is a binary mask or a multi-class mask, which identifies the spatial location and range of the target object in the image information.
[0079] Understandably, spatial feature extraction includes operations such as edge detection and shape descriptor extraction, which are used to capture the geometric features of the target object.
[0080] Understandably, semantic feature extraction uses convolutional neural networks to extract high-level semantic features to capture the semantic attributes of a target object.
[0081] In one implementation, spatial and / or semantic features are extracted from the target object in the image information to obtain the target object's feature information. This method enables direct feature extraction of the target object region in the image without prior image segmentation.
[0082] In one implementation, feature extraction is performed on the image information containing the target object to obtain the image's feature information. For example, the image is encoded using a convolutional neural network to obtain global and / or local feature information. The convolutional neural network can be a backbone network such as ResNet, VGG, or EfficientNet. The global feature information captures the overall semantic and contextual information of the image, such as scene type, lighting conditions, and spatial layout. The local feature information captures local details of the image, such as texture, edges, and color distribution.
[0083] In essence, feature fusion can refer to a mathematical operation and interaction mechanism that integrates features from different information ranges (such as the target object itself and the scene it belongs to) into a unified feature representation. The integration process achieves information interaction, selection, modulation, or juxtaposition through operations between feature elements to generate fused features that possess both target attributes and context awareness, which are then used by downstream derivative prediction models to perform derivative prediction tasks.
[0084] In one implementation, the form of feature fusion includes, but is not limited to, at least one of cross-attention, feature multiplication, and feature splicing.
[0085] In one implementation, cross-attention calculates the correlation between the target object features (including the target object features) and the image features (including the image features) using an attention mechanism. For example, the target object features are used as a query vector, and the image features are used as a key vector and a value vector. The similarity between the query vector and the key vector is calculated using the attention mechanism to obtain attention weights. Then, the value vectors are weighted and summed according to the attention weights to obtain the fused feature information.
[0086] In one implementation, the feature multiplication method multiplies the target object features, which are included in the feature information of the target object, with the image features, which are included in the feature information of the image, element by element. It assumes that the target object features and the image features have a corresponding relationship in the same spatial location, and achieves feature interaction through element-by-element multiplication.
[0087] In one implementation, the feature concatenation method concatenates the target object features (including those in the target object's feature information) with the image features (including those in the image's feature information) along the channel dimension. This method directly concatenates the two types of feature information into a higher-dimensional feature vector, preserving the complete information of both features.
[0088] Through the aforementioned multi-dimensional feature extraction and fusion strategies, the derivative prediction model can comprehensively utilize both the target object's own features and the contextual information of the scene. The target object's own features capture its intrinsic attributes such as shape, edges, and semantics; the contextual information captures external environmental factors such as lighting conditions, scene type, and spatial relationships. The fusion of these two types of feature information enables the derivative prediction model to more accurately infer and predict derivatives associated with the target object.
[0089] For example, in predicting shadow derivatives, the shape features of the target object determine the basic form of the shadow, while the lighting information of the image determines the direction and intensity of the shadow. Through feature fusion, the derivative prediction model can comprehensively consider the target object features and lighting information to accurately predict the spatial location and extent of the shadow.
[0090] For example, in predicting reflection derivatives, the semantic features of the target object determine the content of the reflection, while the spatial layout information of the image determines the position of the reflection. Through feature fusion, the derivative prediction model can comprehensively consider the features of the target object and the spatial layout information to accurately predict the spatial position and range of the reflection.
[0091] The multi-dimensional feature extraction and fusion strategy provided in this embodiment improves the accuracy and robustness of derivative prediction, especially significantly enhancing the prediction performance in complex scenes. Complex scenes include multi-object scenes, occluded scenes, and scenes with varying illumination. In multi-object scenes, contextual information helps distinguish derivatives from different target objects; in occluded scenes, local feature information helps infer derivatives from occluded parts; and in scenes with varying illumination, global illumination information helps predict the impact of illumination changes on derivatives.
[0092] In essence, performing image processing on an object in an image means, based on the region positioning information of the object, performing editing processing on the pixel region covered by the object in the image information to change the visual presentation state of the object.
[0093] In one implementation, image processing includes, but is not limited to, at least one of the following: Editing or processing a portion or the entirety of the object to be processed in a frame of a video; Editing is performed on a portion or the entire object to be processed in an image.
[0094] Intuitively, editing can refer to modifying, replacing, or enhancing the pixel content of an object in an image based on the location of that object, in order to change the object's state of existence or perceptual attributes in visual presentation.
[0095] In one implementation, the editing process includes at least one of elimination, blurring, and image quality optimization.
[0096] The aforementioned technical solution provides a flexible selection of the processing granularity (partial or complete) for the objects to be processed, adapting to different scenario requirements and avoiding over-processing or under-processing. Therefore, the technical solution of this embodiment can achieve precise control of the processing range and realize refined editing.
[0097] In one implementation, the derivative prediction model can be constructed by: constructing multiple training sample data; and training a base model based on the constructed training sample data to construct the derivative prediction model.
[0098] In one implementation, the training sample data includes image information, a target object, feature extraction results of the target object, and corresponding derivative annotations. For example, the target object in the training sample data is determined through manual annotation or automatic segmentation methods; the feature extraction results of the target object in the training sample data can be obtained according to the aforementioned extraction strategy; the derivative annotations in the training sample data can be manually annotated derivative masks that identify the spatial location and range of derivatives associated with the target object.
[0099] In one implementation, the training sample data needs to cover multiple scenarios and various types of derivatives. For example, for shadow derivatives, the training sample data should include shadows under different lighting conditions; for reflection derivatives, the training sample data should include reflections on different reflective surfaces; and for trajectory afterimage derivatives, the training sample data should include afterimages with different motion speeds and trajectories.
[0100] Understandably, the base model can refer to a prototype deep learning network architecture used to build a derivative prediction model. This architecture provides the basic computational capabilities to extract features from image information and establish mapping relationships between features, and its network parameters can be optimized through the training process of training sample data to obtain specialized prediction capabilities for derivative prediction tasks.
[0101] For example, the base model can be constructed based on convolutional networks and / or Transformer networks. Convolutional networks can support derivative inference prediction based on local receptive fields, while Transformer networks can support derivative inference prediction based on global dynamic interactions.
[0102] For example, convolutional networks include, but are not limited to, architectures such as U-Net, DeepLab, and FCN, which are suitable for image segmentation tasks.
[0103] For example, Transformer networks include architectures such as Vision Transformer, Swing Transformer, and SegFormer, which can capture long-distance dependencies and are suitable for derivative prediction in complex scenarios.
[0104] In one implementation, the base model can employ an encoder-decoder architecture. The encoder encodes the feature extraction results, extracting multi-scale features; the decoder decodes the encoded features, progressively restoring the spatial resolution, and finally outputting a derivative mask. Feature fusion between the encoder and decoder can be achieved through skip connections, preserving the detailed information in the encoder.
[0105] In another implementation, the base model can employ a conditional generative model architecture. The feature extraction results are used as conditional inputs, and a derivative mask is generated through a conditional generative network. The conditional generative model can be a conditional GAN, a conditional VAE, a conditional diffusion model, etc.
[0106] In one implementation, a loss function can be configured during model training. The loss function plays a core role in defining the optimization objective, measuring prediction bias, generating gradient signals, balancing multiple task objectives, and driving convergence judgment.
[0107] In one implementation, the loss function during model training is obtained by combining the joint binary cross-entropy loss function and the boundary optimization loss function.
[0108] The binary cross-entropy loss function is used for pixel-level classification of the derivative mask. For each pixel in the derivative mask, the binary cross-entropy loss between the predicted value and the labeled value is calculated. Specifically, the joint binary cross-entropy loss function constrains the classification correctness of the derivative region through pixel-by-pixel probability comparison.
[0109] The boundary optimization loss function is used to refine the precision of the derivative boundaries. Specifically, it calculates the distance error between the predicted mask boundary and the labeled mask boundary. The introduction of the boundary optimization loss function makes the model more focused on the accuracy of the derivative boundaries, avoiding boundary blurring or offset.
[0110] For example, the boundary optimization loss function can be boundary IoU loss, boundary F1 loss, boundary Hausdorff distance loss, etc.
[0111] Among them, the loss function obtained by combining the binary cross-entropy loss function and the boundary optimization loss function can achieve synergistic optimization of regional integrity and boundary clarity through weight adjustment. The loss function transforms the quality requirements of the derivative prediction task into a calculable numerical objective, and transforms the prediction deviation into the update direction of the model parameters through the backpropagation mechanism, ultimately driving the derivative prediction model to converge from the initial parameter state to an optimized state with high-precision regional coverage and boundary positioning capabilities.
[0112] In one implementation, model training can be configured with a learning rate scheduling mechanism. This mechanism plays a crucial role in model training, controlling the parameter update step size, stabilizing the training process, improving convergence accuracy, adapting to multi-task optimization, and protecting pre-trained knowledge. Specifically, the gradual increase in the learning rate during the warm-up phase prevents early gradient instability from damaging pre-trained weights; the high learning rate during the stabilization phase enables rapid exploration of the loss surface and escape from local optima; and the gradual decrease in the learning rate during the decay phase achieves refined searching of the optimal solution basin and suppression of parameter oscillations. Ultimately, this allows the derivative prediction model to achieve efficient and stable convergence in the joint optimization objectives of region classification accuracy and boundary localization precision. The learning rate scheduling mechanism works synergistically with the loss function; the former controls "how fast to go," while the latter defines "where to go," jointly determining the convergence efficiency and final performance of the model training.
[0113] Understandably, a learning rate scheduling mechanism refers to an optimization control strategy that dynamically adjusts the learning rate according to the training rounds or iteration progress and a preset function mapping rule during model training. The optimization control strategy enables the learning rate to transition from a higher initial value to a lower value in the later stages, so as to achieve rapid exploration in the early stages of training and fine convergence in the later stages of training.
[0114] In one implementation, the learning rate scheduling mechanism used for model training includes, but is not limited to, other functional forms with equivalent decay effects such as cosine annealing or linear descent.
[0115] Among them, the cosine annealing mechanism is based on the periodic mapping of trigonometric functions, which causes the learning rate to decrease periodically according to the cosine function.
[0116] The linear descent mechanism is based on a uniform linear function mapping, which causes the learning rate to decrease linearly to a minimum value.
[0117] For example, the specific process of model training is as follows: Training sample data is input into the base model, which predicts a derivative mask based on the feature extraction results of the target object; the loss between the predicted mask and the labeled mask is calculated; model parameters are updated based on backpropagation of the loss; and the learning rate is gradually reduced using a cosine annealing or linear decreasing learning rate scheduling mechanism until the model converges. The criteria for judging model convergence can be that the loss on the validation set no longer decreases, or that the evaluation metrics on the validation set reach a preset threshold. Evaluation metrics include IoU, Dice coefficient, accuracy, and recall.
[0118] The derivative prediction model provided in this embodiment constructs a high-precision derivative prediction model through a deep learning model, a joint loss function, and a learning rate scheduling mechanism, achieving accurate prediction of complex derivatives. These complex derivatives include fluid diffusers and optical phenomena. Fluid diffusers, such as smoke and flames, have irregular boundaries and dynamically changing shapes, making accurate prediction difficult using traditional methods. Optical phenomena, such as shadows and reflections, are affected by lighting conditions and spatial layout, making modeling difficult using traditional methods. This embodiment learns the feature representation of complex derivatives through a deep learning model, optimizes pixel-level classification and boundary fineness through a joint loss function, and improves training efficiency through a learning rate scheduling mechanism, thereby achieving accurate prediction of complex derivatives.
[0119] In one embodiment, the processing method of this embodiment may further include: displaying image information; highlighting the object to be processed in the displayed image information; and adjusting and updating the object to be processed in response to receiving an editing operation on the object to be processed.
[0120] In one implementation, highlighting includes methods such as highlighting, boundary marking, and color overlay. For example, the boundaries of the object to be processed are highlighted, and a striking color is used to indicate the range of the object; or a semi-transparent color overlay is applied to the area of the object to be processed, allowing the user to clearly see the location and range of the object. Through highlighting, users can intuitively view the object to be processed output by the derivative prediction model and judge the accuracy of the derivative prediction.
[0121] In one implementation, editing operations include expanding the scope of the object to be processed, narrowing the scope of the object to be processed, modifying the boundaries of the object to be processed, adding missing derivatives, and deleting erroneous derivatives. In response to receiving an editing operation on the object to be processed, the object to be processed is adjusted and updated. For example, if the user believes the derivative prediction result is inaccurate, the object to be processed can be manually adjusted. Furthermore, if the derivative prediction model misses some shadow areas, the user can add the missing shadow areas to the object to be processed using a selection smear operation. Also, if the derivative prediction model incorrectly identifies a background area as a derivative, the user can delete the erroneous area from the object to be processed using an erase operation.
[0122] Through the aforementioned technical solution, this embodiment can perform secondary adjustments (such as manual editing) on the predicted object to be processed, so that the object to be processed better meets the user's needs and improves the accuracy of image processing.
[0123] In one embodiment, the processing method provided by this embodiment further includes at least one of the following: Based on the objects to be processed before the update, obtain negative training sample data; Based on the updated object to be processed, obtain the positive training sample data; The derivative prediction model is adjusted by processing the model parameters based on negative training sample data and / or positive training sample data.
[0124] Among them, the negative training sample data are samples that the derivative prediction model mispredicted, reflecting the shortcomings of the model.
[0125] Among them, the positive training sample data are the correct samples after user correction, reflecting the user's expected results for derivative prediction.
[0126] For example, both negative training sample data and positive training sample data can include image information, target object, feature extraction results of the target object, and corresponding derivative annotations.
[0127] In one implementation, the derivative prediction model is trained online, incrementally, or fine-tuned using both negative and positive training sample data. Online learning refers to continuously receiving new training sample data and updating the model parameters in real time during model operation. Incremental learning involves training the existing model using new training sample data, retaining previously learned knowledge while learning new knowledge. Fine-tuning involves adjusting the parameters of the existing model using new training sample data to adapt to the new data distribution.
[0128] For example, the specific process of model parameter tuning is as follows: input negative training sample data and positive training sample data into the derivative prediction model, calculate the loss between the prediction result and the positive training sample data, and update the model parameters based on backpropagation of the loss. To avoid catastrophic forgetting, regularization techniques, knowledge distillation techniques, or replay techniques can be used to preserve the model's performance on the original training data.
[0129] Through model parameter tuning, the derivative prediction model can be continuously optimized based on user feedback, gradually improving the accuracy of derivative prediction.
[0130] For example, the interactive feedback process in this embodiment is as follows: the derivative prediction model outputs the object to be processed; the terminal device displays and highlights the object to be processed; the user views and determines whether adjustment is needed; if adjustment is needed, the user performs an editing operation; the terminal device updates the object to be processed according to the editing operation; the objects to be processed before and after the update are collected as negative and positive samples; the positive and negative samples are used for model parameter debugging; and the optimized model is used for subsequent derivative prediction. Through the interactive feedback process provided in this embodiment, continuous optimization and personalized adaptation of the derivative prediction model are achieved. Continuous optimization means that the model can continuously improve based on user feedback, thereby increasing the accuracy of derivative prediction. Personalized adaptation means that the model can adapt to specific scenarios or the habits of specific users. For example, some users tend to retain more derivatives, while some users tend to remove all derivatives. The model can learn user preferences based on historical user feedback and provide personalized derivative prediction results.
[0131] The interactive feedback mechanism enhances the user experience and the naturalness of the processing results. Users are no longer passively receiving derivative prediction results, but can actively participate in the derivative prediction process, adjusting and optimizing the prediction results. At the same time, user feedback data can be used for model optimization, making the model more accurate with use, forming a positive feedback loop.
[0132] Based on the technical concept of the foregoing embodiments, the foregoing embodiments will be illustrated by specific scenario examples below: In one example, a processing method includes: In response to a selection operation on at least a portion of the image information, a target object is determined; Based on the target object and extraction strategy, feature extraction processing is performed to obtain the feature extraction results of the target object; Based on the feature extraction results and derivative prediction model, derivative prediction of the target object is performed to obtain the object to be processed in the image information. The object to be processed includes the target object and derivatives. Perform image processing on the objects to be processed in the image information.
[0133] Specifically, in response to a selection operation on at least a portion of the image information, the target object is determined as follows: the user performs a selection operation (such as clicking, selecting, or smearing) on a specific frame (such as any frame) in the displayed video to determine the area where the target object is located; the specific frame is processed according to a preset processing model and the determined target object to generate an image mask of the target object and a multi-scale semantic feature vector of the entire frame (i.e., the feature information of the image, encoding the hierarchical representation of the image from low-level texture to high-level semantics).
[0134] The preset processing model is an interactive visual understanding model whose functions include: responding to user selection prompts for specific areas of the screen, performing precise segmentation of the target object to generate a pixel-level image mask; and simultaneously extracting multi-scale semantic feature representations of the entire screen to encode hierarchical information from low-level visual attributes to high-level semantic concepts.
[0135] For example, the preset processing model is an interactive segmentation model (such as any SAM series model (such as SAM, SAM2, SAM3), SEEM, LISA, etc.) or a visual-language segmentation model (such as CLIPSeg, LAVT, CRIS, etc.).
[0136] In this example, the preferred pre-configured processing model is the SAM series model. Its interaction methods support various prompts such as points, boxes, and smears. Its segmentation capability supports zero-sample / few-sample accurate segmentation to generate image masks of target objects. It uses a ViT / CNN backbone network to extract multi-scale features and output multi-scale semantic feature vectors of frames. The SAM series model's architecture, which combines a prompt encoder, an image encoder, and a mask decoder, can unify segmentation and feature extraction.
[0137] The process involves feature extraction based on the target object and extraction strategy to obtain the feature extraction results of the target object. Specifically, the obtained image mask of the target object is input into the encoder module. The encoder uses a convolutional neural network (CNN) or Transformer coding structure to extract the spatial and semantic features of the image mask of the target object, resulting in an object feature representation vector (or Mask Embedding, i.e., the aforementioned feature information of the target object). The object feature representation vector and the aforementioned multi-scale semantic feature vector (or Image Embedding) of the entire frame are fused in a feature fusion manner to obtain a fused feature map (i.e., the aforementioned fused feature information) as the feature extraction result, in order to establish a correlation mapping between the target object and its frame context.
[0138] Optionally, the feature fusion method preferably adopts at least one of the following forms: cross-attention, Hadamard product, and concatenation.
[0139] The process involves predicting derivatives of the target object based on feature extraction results and a derivative prediction model to obtain the object to be processed in the image information. The object to be processed includes the target object and its derivatives. Image processing is then performed on the object to be processed in the image information. Specifically, the fused feature map is input into the derivative prediction model to obtain the image mask of the derivative (or Derivative Object Mask). The image mask of the derivative is then logically ORed with the image mask of the target object to obtain the mask of the complete elimination region (i.e., the image mask of the object to be processed). This allows for the execution of an elimination task based on the mask of the complete elimination region. The elimination task can involve feeding a specific frame and the mask of the complete elimination region into the elimination model for elimination processing, and then outputting the eliminated image model for application in a specific frame or multiple frames of the video. Each frame in the multiple frames includes part or all of the object to be processed corresponding to the complete elimination region.
[0140] Among them, the derivative prediction model includes one of the following structural forms (i.e. the aforementioned basic model): (1) Convolutional structure (convolutional network): multi-layer convolutional units, residual blocks and upsampling modules, preferably using a multi-scale decoder combining convolution and deconvolution at the end to preserve local details and capture neighborhood dependencies; (2) Transformer structure (Transformer network): multi-layer self-attention modules combined with a feedforward network to capture global semantic information and long-distance dependencies, and restore spatial resolution in the decoding stage; (3) Hybrid structure: a combination of convolutional structure and Transformer structure, first obtaining local details through convolution, and then using Transformer to capture global features, to achieve more accurate segmentation of derivative regions.
[0141] The training process of the derivative prediction model employs an end-to-end approach to optimize network parameters, specifically including: 1. Training Dataset: A collection of training sample data, containing high-quality labeled data samples of objects and their derivatives, preferably covering various derivative types such as shadows, reflections, transmitted virtual images, specular reflections, and highlights. Training sample data can come from manually labeled data, semi-automatic detection and correction, or publicly available derivative datasets; 2. Input and monitoring signals: A. Input: The image mask of the object is encoded and fused with the multi-scale semantic feature vector of the entire frame to form a feature map, which serves as the network input; B. Supervision signal: The image mask labels of the derivative are used as supervision to perform pixel-by-pixel comparison on the predicted mask output by the network; C. Loss Function: Binary Cross-Entropy Loss and Boundary Loss are used together to improve the segmentation accuracy and edge clarity of the mask; D. Optimization strategy: AdamW or SGD optimizers are preferred, and cosine annealing or linear decreasing learning rate scheduling mechanisms can be adopted.
[0142] Thus, the technical solution in this example can achieve intelligent and rapid mask generation and removal of target objects and their derivatives in a video.
[0143] This application also provides a terminal device, including a memory and a processor. The memory stores a processing program or instructions, which, when executed by the processor, implement the steps of the processing method in any of the above embodiments.
[0144] This application also provides a storage medium storing a processing program or instructions, which, when executed by a processor, implement the steps of the processing method in any of the above embodiments.
[0145] In the embodiments of the terminal device and storage medium provided in this application, all the technical features of any of the above-described processing method embodiments may be included. The extended and explanatory content of the specification is basically the same as that of the embodiments of the above methods, and will not be repeated here.
[0146] This application also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to perform the methods described in the various possible implementations above.
[0147] This application also provides a chip, including a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that a device with the chip installed performs the methods described in the various possible implementations above.
[0148] It is understood that the above scenarios are merely examples and do not constitute a limitation on the application scenarios of the technical solutions provided in the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, as those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0149] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0150] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.
[0151] The units in the device of this application embodiment can be merged, divided, and deleted according to actual needs.
[0152] In this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions are generally described in detail only when they appear for the first time. When they appear again, they are generally not repeated for the sake of brevity. When understanding the technical solutions and other contents of this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions that are not described in detail later can be referred to their previous relevant detailed descriptions.
[0153] In this application, the descriptions of the various embodiments have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0154] The technical features of the present application can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present application.
[0155] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the methods of each embodiment of this application.
[0156] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, storage disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0157] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A processing method, characterized in that, include: Based on the feature extraction results of the target object in the image information and the derivative prediction model, the derivative prediction of the target object is performed to obtain the object to be processed in the image information, the object to be processed including the target object and the derivative; Image processing is performed on the object to be processed in the image information.
2. The processing method as described in claim 1, characterized in that, The image processing includes at least one of the following: Editing is performed on a portion or the entirety of the object to be processed in a frame of the video. Editing or processing a portion or the entirety of the object to be processed in the image; The editing process includes at least one of the following: elimination, blurring, and image quality optimization.
3. The processing method as described in claim 1, characterized in that, The method for determining the target object includes one of the following: In response to a selection operation on at least a portion of a frame of a video, the selected region is designated as the target object; In response to a selection operation on at least a portion of an image frame, the selected area is designated as the target object.
4. The processing method according to any one of claims 1 to 3, characterized in that, The extraction strategy for the feature extraction results includes at least one of the following: Feature extraction is performed on the target object to obtain its feature information; Feature extraction is performed on the image information containing the target object to obtain the feature information of the image; The feature information of the target object and the feature information of the image are fused to obtain the fused feature information; The feature extraction result is determined based on the feature information of the target object or the fused feature information; The feature fusion forms include at least one of cross-attention, feature multiplication, and feature splicing.
5. The processing method as described in claim 4, characterized in that, The feature extraction of the target object includes at least one of the following: Based on the target object and the image segmentation model, object segmentation is performed to obtain the image mask of the target object; Spatial and / or semantic features are extracted from the image mask of the target object to obtain the feature information of the target object; Spatial feature extraction and / or semantic feature extraction are performed on the target object in the image information to obtain the feature information of the target object.
6. The processing method according to any one of claims 1 to 3, characterized in that, The derivative prediction model is constructed in a manner that includes at least one of the following: Construct multiple training sample data; The base model is trained based on the constructed training sample data to construct the derivative prediction model; The basic model is constructed based on convolutional networks and / or Transformer networks; The loss function used in the model training is obtained by combining the joint binary cross-entropy loss function and the boundary optimization loss function. The learning rate scheduling mechanism used in the model training includes cosine annealing or linear descent.
7. The processing method according to any one of claims 1 to 3, characterized in that, The processing method further includes: Display the image information; The object to be processed is highlighted in the displayed image information; In response to receiving an edit operation on the object to be processed, the object to be processed is adjusted and updated.
8. The processing method as described in claim 7, characterized in that, The processing method further includes at least one of the following: Based on the object to be processed before the update, obtain negative training sample data; Based on the updated object to be processed, obtain positive training sample data; The derivative prediction model is adjusted based on the negative training sample data and / or the positive training sample data.
9. A terminal device, characterized in that, include: A memory and a processor, the memory storing a processing program or instructions which, when executed by the processor, implement the processing method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, The storage medium stores a computer program or instructions, which, when executed by a terminal device, implement the processing method as described in any one of claims 1 to 8.