Multimedia data processing method and apparatus
By employing targeted optimization strategies after AI model processing to eliminate quality defects in multimedia data, the new defects introduced by AI models are resolved, thereby improving the quality of multimedia data and user experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-04-02
AI Technical Summary
Existing AI models suffer from generalization instability and black-box characteristics when improving the quality of multimedia data. This leads to new defects in the processed video, and eliminating these defects requires a large amount of sample data for adjustment, making the process uncontrollable.
After processing multimedia data using the first optimization strategy, the second optimization strategy is used to specifically eliminate defects within the quality defect area. The defect area is determined by combining user interaction or automatic identification methods, and optimization is performed using AI or non-AI models to ensure accuracy and a low probability of introducing new defects.
It effectively eliminates quality defects in multimedia data, improves processing quality, reduces user operation difficulty and time, and enhances user experience.
Smart Images

Figure CN2025123917_02042026_PF_FP_ABST
Abstract
Description
Method and apparatus for processing multimedia data
[0001] The present application claims priority to the Chinese patent application No. 202411379911.X, filed on September 29, 2024, and entitled "Method and apparatus for processing multimedia data", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of multimedia processing, and in particular, to a method and apparatus for processing multimedia data. BACKGROUND
[0003] With the continuous development of artificial intelligence (AI) technology, the functions of AI models are becoming more and more powerful.
[0004] Currently, AI models can be applied in quality improvement of multimedia data (images, audio, video streams or other forms). For example, using AI models to perform noise reduction, resolution adjustment, flicker removal, color enhancement, etc. on the captured video greatly improves the image quality in the video stream.
[0005] However, due to the instability of generalization and the black box characteristics of AI models, while the AI model improves the video quality, the model itself causes new defects (defects are used to describe quality problems of data) in the processed video. Eliminating defects requires collecting a large amount of sample data to adjust the model, and the process is not controllable. Therefore, optimizing the defects introduced by AI models in processing multimedia data has become the focus of industry research. SUMMARY
[0006] The present application provides a method and apparatus for processing multimedia data to optimize the defects introduced by the strategy in processing multimedia data and improve the quality of the strategy in processing multimedia data.
[0007] To achieve the above-mentioned purpose, the present application adopts the following technical solutions.
[0008] In a first aspect, a method for processing multimedia data is provided, comprising: determining at least one quality defect area of a target data unit existing in multimedia data processed by a first optimization strategy. Based on the at least one quality defect area, determining at least one second optimization strategy. Based on the at least one second optimization strategy, optimizing the target data unit. Wherein, the target data unit is one or more data units in the multimedia data, and the multimedia data includes any of the following types: image, video, or audio.
[0009] According to the scheme provided in the application, after the multimedia data is processed by the first optimization strategy, the target data unit is optimized by using the second optimization strategy, and the quality defects in the quality defect area of the target data unit are eliminated. The second optimization strategy for eliminating quality defects can be configured according to actual needs, and the accuracy of eliminating quality defects by using the second optimization strategy is ensured. In this way, even if the multimedia data processed by the first optimization strategy introduces quality defects, the second optimization strategy can be used to supplement the quality defect area (part of the multimedia data) where the quality defects exist, so as to eliminate the quality defects in the quality defect area. In addition, the second optimization strategy only processes the quality defect area, and the probability of introducing new quality defects is very low. Therefore, the scheme provided in the application optimizes the defects introduced by the strategy processing the multimedia data, and improves the quality of the strategy processing the multimedia data.
[0010] In a possible implementation, one second optimization strategy is used to optimize one type of quality defects associated therewith. The above determining at least one second optimization strategy based on at least one quality defect area can be specifically implemented as: determining at least one second optimization strategy associated with at least one type of quality defects existing in at least one quality defect area. Through the pre-configured association between the type of quality defects and the second optimization strategy, the second optimization strategy is only used to optimize one type of quality defects associated therewith, which can ensure the accuracy of the second optimization strategy in optimizing the quality defects and reduce the probability of introducing new quality defects. In this way, the second optimization strategy is used to optimize the quality defect area, which can accurately optimize the quality defects therein and almost does not introduce new quality defects.
[0011] In another possible implementation, the above determining at least one quality defect area existing in the target data unit in the multimedia data processed by the first optimization strategy can be specifically implemented as: receiving a first user instruction for indicating one or more quality defect areas, and determining the quality defect area indicated by the first user instruction as at least one quality defect area existing in the target data unit. Through the user interaction, the user selects the target data unit that needs to be optimized for quality defects, so that the optimization of the target data unit meets the user's needs and improves the user experience.
[0012] In another possible implementation, different quality defect types correspond to different matching features, and the determining the at least one quality defect region existing in the target data unit in the multimedia data processed by using the first optimization strategy can be specifically implemented as follows: receiving a first user instruction for indicating one or more quality defect regions, identifying other regions in the target data unit except for the quality defect region indicated by the first user instruction based on a matching feature corresponding to at least one quality defect type existing in the quality defect region indicated by the first user instruction, and obtaining a quality defect region in which at least one quality defect of the quality defect type exists in the other regions. The quality defect region indicated by the first user instruction and the quality defect region in the other regions are determined as the at least one quality defect region existing in the target data unit.
[0013] By means of user interaction, a user selects a partial quality defect region in the target data unit that needs to be optimized, and then selects other quality defect regions having the same quality defect type as the partial quality defect region in other positions in the target data unit in a feature matching manner. In this way, the quality defect region of the target data unit is quickly determined by combining user interaction and machine intelligent identification, the optimization time of the quality defect is shortened, and the automation degree and quality of the optimization of the quality defect are improved. In addition, the user does not need to select all the quality defect regions, the operation difficulty of the user is reduced, and the experience of the user is improved.
[0014] In another possible implementation, the determining the at least one quality defect region existing in the target data unit in the multimedia data processed by using the first optimization strategy can be specifically implemented as follows: identifying the target data unit by using a first identification strategy to obtain the at least one quality defect region. The quality defect region is automatically identified and obtained in the target data unit by using the first identification strategy, which can more intelligently and efficiently obtain the quality defect region, and make the process of eliminating the quality defect intelligent and efficient.
[0015] In another possible implementation, the method for processing multimedia data provided in the present application further includes: identifying the at least one quality defect region by using a second identification strategy to obtain at least one quality defect type existing in the at least one quality defect region. In the process of processing the multimedia data, the quality defect types in different quality defect regions can be adaptively determined by using the second identification strategy, and the acquisition efficiency of the quality defect types is improved.
[0016] In a possible implementation, the method for processing multimedia data further includes: receiving a second user instruction indicating a quality defect type. The quality defect type indicated by the second user instruction is determined as at least one quality defect type existing in the at least one quality defect area. The quality defect type in the quality defect area indicated by the user is beneficial to eliminate the quality defect of the specified type, so that the processed multimedia data meets the user expectation.
[0017] In a possible implementation, the method for processing multimedia data further includes: fusing the target data unit optimized by the second optimization strategy into the multimedia data according to the position of the target data unit in the multimedia data. In this way, the purpose of optimizing the multimedia data by the second optimization strategy is achieved, and the quality of the multimedia data is improved.
[0018] In a possible implementation, the second optimization strategy includes any one of the following: an AI model, a non-AI model for low-frequency filtering, a non-AI model for edge detection, or a non-AI model for color correction. In this way, the quality defect in the multimedia data is optimized by the AI model or the non-AI model, and the universality of the scheme is enhanced.
[0019] In a possible implementation, the quality defect type includes any one of the following: detail smearing loss, texture performance inter-frame inconsistency, noise flicker, bright noise, color noise, color edge, contour convexity defect, color oversaturation, ghosting, pseudo color, stroboscopic, or color cast. The multiple types of quality defects in the multimedia data can be coped with and optimized, so that the application scenarios of the scheme are more extensive.
[0020] In a second aspect, a device for processing multimedia data is provided. The device can include a first determining module, a second determining module, and an optimization module. Wherein:
[0021] The first determining module is configured to determine at least one quality defect area existing in a target data unit in the multimedia data processed by a first optimization strategy. The target data unit is one or more data units in the multimedia data. The multimedia data includes any one of the following types: an image, a video, or an audio.
[0022] The second determining module is configured to determine at least one second optimization strategy based on the at least one quality defect area.
[0023] The optimization module is configured to optimize the target data unit based on the at least one second optimization strategy.
[0024] In a possible implementation, one second optimization strategy is used to optimize one type of quality defect associated therewith, and the second determining module is specifically configured to determine at least one second optimization strategy associated with at least one type of quality defect existing in at least one quality defect region. Through the preconfigured association between the type of quality defect and the second optimization strategy, the second optimization strategy is used to optimize only one type of quality defect associated therewith, so as to ensure the accuracy of the second optimization strategy in optimizing the quality defect and reduce the probability of introducing new quality defects. In this way, the second optimization strategy is used to optimize the quality defect region, so that the quality defect in the quality defect region can be accurately optimized, and almost no new quality defect is introduced.
[0025] In another possible implementation, the first determining module is specifically configured to: receive a first user instruction used to indicate one or more quality defect regions; and determine the quality defect region indicated by the first user instruction as the at least one quality defect region existing in the target data unit. Through user interaction, the user selects the target data unit that needs to be optimized in terms of quality defect, so that the optimization of the target data unit meets the user demand, and the user experience is improved.
[0026] In another possible implementation, different types of quality defects correspond to different matching features, and the first determining module is specifically configured to: receive a first user instruction used to indicate one or more quality defect regions; identify other regions in the target data unit other than the quality defect region indicated by the first user instruction, based on matching features corresponding to at least one type of quality defect existing in the quality defect region indicated by the first user instruction, to obtain a quality defect region in which at least one type of quality defect existing in the other regions is indicated; and determine the quality defect region indicated by the first user instruction and the quality defect region in the other regions as the at least one quality defect region existing in the target data unit. Through user interaction, the user selects part of the quality defect region in the target data unit that needs to be optimized in terms of quality defect, and then screens other quality defect regions having the same type of quality defect as the quality defect region in the part of the quality defect region in the target data unit in a feature matching manner. In this way, the quality defect region of the target data unit is quickly determined through the combination of user interaction and machine intelligent identification, the optimization time of the quality defect is shortened, and the automation degree and quality of the optimization of the quality defect are improved. In addition, the user does not need to select all the quality defect regions, the operation difficulty of the user is reduced, and the user experience is improved.
[0027] In a possible implementation, the first determining module is specifically configured to: adopt a first identification strategy to identify the target data unit to obtain at least one quality defect area in the target data unit. By adopting the first identification strategy to automatically identify the quality defect area in the target data unit, the quality defect area can be more intelligently and efficiently obtained, and the process of eliminating the quality defect is intelligent and efficient.
[0028] In a possible implementation, the apparatus further includes an identification module configured to adopt a second identification strategy to identify the at least one quality defect area to obtain at least one quality defect type existing in the at least one quality defect area. In the process of processing the multimedia data, the quality defect type in different quality defect areas can be adaptively determined by adopting the second identification strategy, and the efficiency of obtaining the quality defect type is improved.
[0029] In a possible implementation, the apparatus further includes a second receiving unit and a second determining unit. The second receiving unit is configured to receive a second user instruction used to indicate the quality defect type. The second determining unit is configured to determine the quality defect type indicated by the second user instruction as the at least one quality defect type existing in the at least one quality defect area. The quality defect type in the quality defect area is indicated by the user, which is beneficial to eliminating the quality defect of the specified type, and the processed multimedia data meets the user's expectation.
[0030] In a possible implementation, the apparatus further includes a fusion module configured to fuse the target data unit optimized by the second optimization strategy to the multimedia data according to the position of the target data unit in the multimedia data. In this way, the purpose of optimizing the multimedia data by adopting the second optimization strategy is achieved, and the quality of the multimedia data is improved.
[0031] In a possible implementation, the second optimization strategy includes any one of the following: an AI model, a non-AI model used for low-frequency filtering, a non-AI model used for edge detection, or a non-AI model used for color correction. In this way, the quality defect in the multimedia data is optimized by the AI model or the non-AI model, and the universality of the scheme is enhanced.
[0032] In a possible implementation, the quality defect type includes any one of the following: loss of detail smearing, inter-frame inconsistency of texture performance, noise flicker, bright noise, color noise, color edge, contour convexity defect, color oversaturation, ghosting, pseudo color, stroboscopic, or color cast. The multiple types of quality defects in the multimedia data can be coped with and optimized, and the application scenarios of the scheme are more extensive.
[0033] In a third aspect, a server is provided, comprising a processor and a memory, the memory being configured to store computer instructions, which, when executed by the processor, cause the server to perform the method in the first aspect or any possible implementation thereof.
[0034] In a fourth aspect, an electronic device is provided, comprising a processor and a memory, the memory being configured to store computer instructions, which, when executed by the processor, cause the electronic device to perform the method in the first aspect or any possible implementation thereof.
[0035] In a fifth aspect, a chip is provided, comprising: a processor, the processor being configured to execute computer instructions to perform the operation steps of the method in the first aspect or any possible implementation thereof.
[0036] In a possible implementation, the chip described in the fourth aspect further comprises a memory, configured to store computer instructions.
[0037] In a sixth aspect, a computer program product is provided, comprising computer programs or instructions, which, when executed on a processor, cause the processor to perform the operation steps of the method in the first aspect or any possible implementation of the first aspect.
[0038] In a seventh aspect, a computer-readable storage medium is provided, comprising: computer instructions or programs; when the computer instructions or programs are executed on a computer, the computer executes the operation steps of the method in the first aspect or any possible implementation of the first aspect.
[0039] The technical effects brought by any implementation of the third aspect to the seventh aspect can refer to the technical effects brought by the corresponding implementation of the first aspect, which will not be repeated here.
[0040] It should be noted that the various possible implementations of any one of the above aspects can be combined as long as the schemes are not contradictory. BRIEF DESCRIPTION OF DRAWINGS
[0041] FIG. 1 is a flowchart of a processing method of a video denoising model;
[0042] FIG. 2 is a schematic diagram of a principle architecture of a video enhancement scheme;
[0043] FIG. 3 is a flowchart of an erasing method for eliminating redundant parts of a picture;
[0044] FIG. 4 is a schematic diagram of a system for processing multimedia data provided by the present application;
[0045] FIG. 5 is a schematic diagram of a structure of a terminal device provided by the present application;
[0046] FIG. 6 is a flowchart of a method for processing multimedia data according to an embodiment of the present application;
[0047] FIG. 7 is a flowchart of another method for processing multimedia data according to an embodiment of the present application;
[0048] FIG. 8 is a schematic diagram of a system architecture for processing multimedia data according to an embodiment of the present application;
[0049] FIG. 9 is a schematic diagram of another system architecture for processing multimedia data according to an embodiment of the present application;
[0050] FIG. 10 is a schematic diagram of an interface of a terminal according to an embodiment of the present application;
[0051] FIG. 11 is a schematic diagram of another interface of a terminal according to an embodiment of the present application;
[0052] FIG. 12 is a schematic diagram of a further interface of a terminal according to an embodiment of the present application;
[0053] FIG. 13 is a schematic diagram of a device for processing multimedia data according to an embodiment of the present application;
[0054] FIG. 14 is a schematic diagram of a server according to an embodiment of the present application. DETAILED DESCRIPTION
[0055] For the convenience of understanding, the main terms involved in the present application are first explained.
[0056] Artificial intelligence (AI) is a new technical science that studies and develops theories, methods, techniques and application systems for simulating, extending and expanding human intelligence. It is a branch of computer science.
[0057] Multimedia data refers to data composed of forms of data such as graphics, images, audio, video and the like that are related to each other in content. The multimedia data can include any one of images, video (which can also be referred to as a video stream, or a sequence of image frames), or audio.
[0058] A data unit refers to an operable unit in multimedia data. The size of the data unit in multimedia data can be configured according to actual needs. For example, when the multimedia data is an image, the data unit is the entire image or an image block; when the multimedia data is a video, the data unit can be a single image frame; and when the multimedia data is audio, the data unit can be an audio segment of a predetermined size.
[0059] A neural network is a machine learning model, which is a kind of neural network simulating human brain to realize machine learning technology of artificial intelligence. The input and output of the neural network can be configured according to actual needs, and the neural network is trained through sample data to minimize the error between the output of the neural network and the real output corresponding to the sample data. The neural network can include but is not limited to a deep neural network, a convolutional neural network, a recurrent neural network, or others.
[0060] Currently, improving the quality of multimedia data optimization of AI models is the focus of industry research, and many solutions have been proposed, which are described below.
[0061] One solution is a video denoising model processing method. FIG. 1 is a flowchart of the video denoising model processing method. As shown in FIG. 1, in this solution, a target video frame is first obtained in a sample video frame sequence, and then an image detail feature of the target video frame is extracted through a first branch of the video denoising model. After obtaining a down-sampled video frame sequence corresponding to the video frame sequence, a feature extraction is performed on the down-sampled video frame sequence through a second branch of the video denoising model to obtain an image fusion feature. Then, a predicted video frame is generated based on the image fusion feature and the image detail feature. Finally, parameters in the video denoising model are adjusted according to a loss value between the predicted video frame and a video frame corresponding to the target video frame in a reference video, and a target video denoising model is obtained. This solution fuses video features through two branches to improve the denoising effect of the video denoising model, but the generalization instability and black box characteristics of the model cannot be avoided. Even if the denoising effect is improved, quality defects caused by the denoising process cannot be avoided. Targeted training of the model may encounter various special scene defects that offset each other.
[0062] Another solution is a video enhancement solution. FIG. 2 is a principle architecture diagram of the video enhancement solution. As shown in FIG. 2, a to-be-processed video is input into a dynamic analysis module. The dynamic analysis module generates a recommended video enhancement operation (super resolution, video denoising, high dynamic range, anti-flicker, slow motion, or color enhancement) according to video corresponding features and in combination with user preferences. An adaptive processing core module performs an enhancement operation on the to-be-processed video according to the recommended video enhancement operation to obtain an enhanced video. The user can feed back the like or dislike of the recommended operation as the user preference to the dynamic analysis module according to the effect of the enhanced video. Due to the unexplainability of machine learning, processing of the user preference can introduce additional image defects, resulting in that the solution cannot obtain feedback of local defects of the image from the user and cannot eliminate the defects in a targeted manner. Quality defects still exist in the enhanced video.
[0063] Another solution is an erasing method for eliminating the redundant part of the picture. FIG. 3 is a flowchart of the erasing method for eliminating the redundant part of the picture. As shown in FIG. 3, in this solution, a picture is first obtained, the picture is preprocessed, the smearing part of the picture is extracted, the target range to be erased is selected, and a sample picture is obtained. The position where the target range appears in the sample picture is located, the histogram of oriented gradient (HOG) feature points and color features that are most matched with the target range are extracted, and the position is marked as a target image to be erased; the marked target image is erased, and the target range is reasonably filled to obtain an effect picture. In this solution, the selected region is matched with features, the erasing object in the region is selected, the target image is erased, and a reasonable image is filled. This method only acts on the selected region, but cannot avoid the quality defects introduced during erasing and filling.
[0064] As can be seen, in the current solution for reducing the defects introduced by the AI model in processing multimedia data, if there is content in the multimedia data that has not been processed by the AI model, or the AI model cannot accurately understand complex task requirements involving emotions, ethics, etc., or due to the opaque decision-making process of the AI model, etc., the processing result of the multimedia data by the AI model will introduce new defects. In order to eliminate these defects, a large amount of sample data can be collected to train the model, but the output result of the adjusted model can be uncontrollable, which leads to new defects in the process of optimizing the defects introduced by the AI model itself. Therefore, the quality of the multimedia data after the optimization of the current AI model is not high.
[0065] Based on this, the present application provides a method for processing multimedia data. After the multimedia data is processed by a first optimization strategy, a second optimization strategy is used to optimize the target data unit, and the quality defects in the quality defect area of the target data unit are eliminated. The second optimization strategy for solving quality defects can be configured according to actual needs to ensure the accuracy of eliminating quality defects by the second optimization strategy. In this way, even if the multimedia data processed by the first optimization strategy introduces quality defects, the second optimization strategy can be used to process only the quality defect area (part of the multimedia data) where the quality defects exist to eliminate the quality defects in the quality defect area; and the second optimization strategy only processes the quality defect area, and the probability of introducing new quality defects is very low. Therefore, the solution of the present application improves the quality of the multimedia data.
[0066] The embodiments of the present application will be described below with reference to the accompanying drawings.
[0067] FIG. 4 is a structural schematic diagram of a system for processing multimedia data according to an embodiment of the present application. As shown in FIG. 4, the system for processing multimedia data includes an electronic device 401 and a computing device 402. The electronic device 401 and the computing device 402 are communicatively connected.
[0068] The electronic device 401 is configured to obtain multimedia data processed by a first optimization strategy, and send the multimedia data to the computing device 402, so that the computing device 402 processes the multimedia data according to the scheme provided in the present application.
[0069] In an embodiment of the present application, the electronic device 401 can generate (capture, record or in other ways) the multimedia data, and optimize the multimedia data by using the first optimization strategy. Alternatively, the electronic device 401 can obtain multimedia data imported by a user, and optimize the multimedia data by using the first optimization strategy. Alternatively, the electronic device 401 can obtain multimedia data processed by using the first optimization strategy from other devices (the computing device 402 or other computing devices or other electronic devices). The present application does not limit the way in which the electronic device 401 obtains the multimedia data processed by using the first optimization strategy.
[0070] For example, the first optimization strategy can be an AI model, or other models or algorithms, which is configured to improve the quality of the multimedia data. The present application does not limit the content of the first optimization strategy.
[0071] For example, the electronic device 401 can be a terminal device, such as a smartphone, a wearable device, an augmented reality (AR) / virtual reality (VR) device, a tablet computer, a notebook computer, a vehicle-mounted device, an ultra-mobile personal computer (UMPC), a netbook, a personal computer (PC) or a personal digital assistant (PDA), etc. The present application does not limit the specific form of the electronic device 401.
[0072] For example, the electronic device 401 can implement the above functions in the form of an application (App).
[0073] In a possible implementation, the computing device 402 can be deployed inside the electronic device 401, and the two devices can be connected by hardware circuits inside the electronic device 401.
[0074] In another possible implementation, the computing device 402 can be deployed outside the electronic device 401, for example, the computing device 402 can be implemented in the form of a cloud service and deployed on a cloud. The electronic device 401 and the computing device 402 can be connected through a communication network.
[0075] The embodiment of the present application does not specifically limit the deployment relationship and the connection relationship of the electronic device 401 and the computing device 402.
[0076] The internal architecture of the electronic device 401 will be exemplarily described below taking the electronic device 401 as a terminal device. Of course, the internal architecture of the electronic device 401 can be configured according to actual needs, and the internal architecture in FIG. 5 is only exemplarily described and does not constitute a specific limitation.
[0077] Please refer to FIG. 5, which is a structural schematic diagram of a terminal device provided by the present application. As shown in FIG. 5, the terminal device includes a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a power management module 140, an antenna, a wireless communication module 160, an audio module 170, a loudspeaker 170A, a sound box interface 170B, a microphone 170C, a sensor module 180, a key 190, an indicator 191, a display screen 192, a camera 193, and the like. The sensor module 180 can include a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, and the like.
[0078] The structure exemplified by the embodiment does not constitute a specific limitation on the terminal device. In other embodiments, the terminal device can include more or fewer components than those shown, or combine certain components, or split certain components, or different arrangement of components. The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0079] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), and the like. Different processing units can be independent devices or integrated in one or more processors.
[0080] The processor 110 can also include memory that stores instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The cache memory can hold instructions or data that the processor 110 has recently used or is likely to use again. If the processor 110 needs to use that instruction or data again, it can be called directly from the cache memory. This avoids repeated accesses and reduces the latency of the processor 110, thus improving the efficiency of the system.
[0081] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, and / or a USB interface, etc.
[0082] The power management module 140 is used to connect a power source. The power management module 140 can also be connected to the processor 110, the internal memory 121, the display screen 192, the camera 193, and the wireless communication module 160, etc. The power management module 140 receives the input of the power source and supplies power to the processor 110, the internal memory 121, the display screen 192, the camera 193, and the wireless communication module 160, etc. In some embodiments, the power management module 140 can also be disposed in the processor 110.
[0083] The wireless communication function of the terminal device can be implemented by the antenna and the wireless communication module 160, etc. The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as Wi-Fi network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied to the terminal device.
[0084] The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive signals to be transmitted from the processor 110, frequency modulate them, amplify them, and radiate them as electromagnetic waves via the antenna. In some embodiments, the antenna and the wireless communication module 160 of the terminal device are coupled, so that the terminal device can communicate with the network and other devices through wireless communication technology.
[0085] In the embodiments of the present application, the terminal device can communicate with the computing device 103 through the wireless communication module 160 and the antenna, and realize the processing of multimedia data.
[0086] The terminal device realizes the display function through the GPU, the display screen 192, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 192 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.
[0087] The display screen 192 is configured to display text, images, and videos, etc. The display screen 192 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diodes (QLED), etc.
[0088] In this embodiment, the display screen 192 is configured to display interfaces provided by the terminal device. The interfaces can be configured to display multimedia data (e.g., images captured by the camera 193, or videos / audio recorded), or display content processed by the computing device 103, etc.
[0089] The terminal device can implement the photographing function through an ISP, the camera 193, a video codec, a GPU, the display screen 192, and an application processor, etc. The ISP is configured to process data fed back by the camera 193. In some embodiments, the ISP can be arranged in the camera 193.
[0090] The camera 193 is configured to capture still images or videos. An object generates an optical image through a lens and projects the optical image to a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert the electrical signal into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format, such as RGB, YUV, etc. In some embodiments, the terminal device can include one or N cameras 193, where N is a positive integer greater than 1.
[0091] Alternatively, the terminal device can not include the camera, i.e., the camera 193 is not arranged in the terminal device (e.g., a television). The terminal device can externally connect the camera 193 through an interface (e.g., the USB interface 130). The externally connected camera 193 can be fixed on the terminal device through an external fixing member (e.g., a camera holder with a clamp). For example, the externally connected camera 193 can be fixed at an edge (e.g., an upper edge) of the display screen 192 of the terminal device through the external fixing member.
[0092] The NPU is a neural-network (NN) calculation processor. By drawing lessons from the structure of a biological neural network, for example, by drawing lessons from the transmission mode between human brain neurons, the NPU can quickly process input information and can continuously self-learn. Through the NPU, intelligent cognition of the terminal device and other applications can be realized, for example, image recognition, face recognition, voice recognition, text understanding, and the like.
[0093] The external memory interface 120 can be used to connect an external memory card, for example, a Micro SD card, to realize the expansion of the storage capacity of the terminal device. The external memory card communicates with the processor 110 through the external memory interface 120 to realize a data storage function. For example, files such as text, images, and videos are saved in the external memory card.
[0094] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various function applications and data processing of the terminal device by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, and the like), and the like. The data storage area can store data (such as audio data and the like) created during the use of the terminal device, and the like. In addition, the internal memory 121 can include a high-speed random access memory and can also include a non-volatile memory, for example, at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like.
[0095] The terminal device can realize audio functions such as music playing and recording through the audio module 170, the speaker 170A, the microphone 170C, the audio interface 170B, the application processor, and the like. In this application, the microphone 170C can be used to receive voice instructions issued by the user to the terminal device. The speaker 170A can be used to feed back decision instructions of the terminal device to the user.
[0096] The audio module 170 is configured to convert digital audio information into an analog audio signal output, and to convert an analog audio input into a digital audio signal. The audio module 170 can also be configured to encode and decode audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some of the functions of the audio module 170 can be disposed in the processor 110. The speaker 170A, also referred to as a "loudspeaker", is configured to convert an audio electrical signal into a sound signal. The microphone 170C, also referred to as a "microphone", "sound transducer", is configured to convert a sound signal into an electrical signal.
[0097] The audio module 170 is configured to convert digital audio information into an analog audio signal output, and to convert an analog audio input into a digital audio signal. The audio module 170 can also be configured to encode and decode audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some of the functions of the audio module 170 can be disposed in the processor 110. The speaker 170A, also referred to as a "loudspeaker", is configured to convert an audio electrical signal into a sound signal. The microphone 170C, also referred to as a "microphone", "sound transducer", is configured to convert a sound signal into an electrical signal.
[0098] The keys 190 include a power key, a volume key, and the like. The keys 190 can be mechanical keys. Alternatively, the keys 190 can be touch keys. The terminal device can receive key input and generate key signal input related to user settings and function control of the terminal device.
[0099] The indicator 191 can be an indicator light, which can be used to indicate that the terminal device is in a powered-on state, a standby state, or a powered-off state, and the like. For example, when the indicator light is off, it can indicate that the terminal device is in a powered-off state; when the indicator light is green or blue, it can indicate that the terminal device is in a powered-on state; and when the indicator light is red, it can indicate that the terminal device is in a standby state.
[0100] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the terminal device. The terminal device can have more or fewer components than those shown in FIG. 5, can combine two or more components, or can have a different component configuration. For example, the terminal device can also include an audio module, and the like. The various components shown in FIG. 5 can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing or application-specific integrated circuits.
[0101] The method for processing multimedia data provided by the present application will be described in detail below with reference to the accompanying drawings.
[0102] In one aspect, FIG. 6 is a flowchart illustrating a method for processing multimedia data, which can be performed by a computing device, such as the computing device 402 of FIG. 4. The computing device can be in communication with an electronic device, such as the electronic device 401 of FIG. 4.
[0103] In particular, after obtaining the multimedia data processed by the first optimization strategy, the electronic device can perform the following method for processing multimedia data by the computing device to optimize the multimedia data.
[0104] The first optimization strategy is used to optimize the multimedia data and improve the quality of the multimedia data. For example, the first optimization strategy is used to eliminate the quality defects associated with the first optimization strategy. When the multimedia data is processed by the first optimization strategy, the quality defects associated with the first optimization strategy in the multimedia data are eliminated.
[0105] The multimedia data can include any of the following types: image, video, or audio.
[0106] The quality defect is the content that deviates from the reference data to the current data to different degrees. The reference data can be data without quality defects, or original data, or data before processing the current data. Alternatively, in the present application, the reference data is the multimedia data processed by the first optimization strategy, and the current data is the multimedia data processed by the first optimization strategy. According to the characteristics of the quality defects, the quality defects can be divided into different quality defect types, including but not limited to any of the following: detail smearing loss, texture performance inconsistency between frames, noise flicker, bright noise, color noise, color edge, contour convexity defect, color oversaturation, ghosting, pseudo color, stroboscopic, or color cast.
[0107] Alternatively, the first optimization strategy is an AI model or a non-AI model strategy.
[0108] Due to the generalization instability and black box characteristics of the model, the multimedia data processed by the first optimization strategy introduces new quality defects, resulting in low quality of the multimedia data. The method for processing multimedia data provided by the present application further optimizes the multimedia data processed by the first optimization strategy to improve the quality of the multimedia data.
[0109] As shown in FIG. 6, the method for processing multimedia data provided by the present application includes:
[0110] S601, the computing device determines at least one quality defect area existing in a target data unit in the multimedia data processed by the first optimization strategy.
[0111] The target data unit is an optimization object when the scheme of the application is used to optimize multimedia data. The number of target data units and the position of the target data units in the multimedia data can be configured according to actual needs, and the embodiments of the application are not limited in this regard. For example, the target data unit is part or all of the data units in the multimedia data.
[0112] In a possible implementation, the target data unit is specified by a user.
[0113] In another possible implementation, the data unit at a preset position in the multimedia data is configured as the target data unit.
[0114] For example, the preset position can be all data units in the multimedia data. Alternatively, the preset position can be the Xth data unit in the multimedia data, and X can be configured according to actual needs, and X can be one or more values. Alternatively, the preset position is the data unit selected by the user, and the preset number of data units. Of course, the content of the preset position can be configured according to actual needs, and the embodiments of the application are not limited in this regard.
[0115] Specifically, the quality defect area is an area in the target data unit where the quality defect to be optimized exists.
[0116] The following examples illustrate the implementation of several target data units and the scheme for determining at least one quality defect area in the target data unit in various implementations, but do not constitute a limitation on the scheme of the application:
[0117] Implementation 1: The quality defect area in the target data unit where the quality defect exists is determined by interacting with the user.
[0118] Specifically, in implementation 1, the multimedia data can be displayed in a preview manner on a human-computer interaction interface (deployed on an electronic device or a computing device), and the user can operate the previewed multimedia data on the human-computer interaction interface.
[0119] In implementation 1, determining the quality defect area in the target data unit where the quality defect exists can include but is not limited to the following cases:
[0120] Case 1: The target data unit and the quality defect area are both specified by the user.
[0121] In case 1, the data unit operated by the user on the human-computer interaction interface is taken as the target data unit, and the quality defect area selected by the user is taken as the quality defect area in the target data unit where the quality defect exists.
[0122] For example, the user can select the quality defect area to be optimized by selecting operation in the human-computer interaction interface. S601 can be specifically implemented as: receiving a first user instruction, the first user instruction being used to indicate one or more quality defect areas, determining the quality defect area indicated by the first user instruction as at least one quality defect area existing in the target data unit, and taking the data unit operated by the user as the target data unit.
[0123] Optionally, the first user instruction can be touch operation, key-based operation, voice instruction or shortcut gesture in the human-computer interaction interface of the electronic device, and the embodiment of the present application is not limited to the implementation form of the first user instruction.
[0124] For example, the selecting operation can be frame selection, circle selection or other operations, and the embodiment of the present application is not limited thereto.
[0125] Case two: the user specifies part of the quality defect area.
[0126] In one possible implementation, in case two, the data unit operated by the user in the human-computer interaction interface is taken as the target data unit.
[0127] In another possible implementation, in case two, the target data unit is the data unit at a preset position in the multimedia data.
[0128] In case two, according to the matching feature of the quality defect type in the quality defect area selected by the user, the quality defect area in other areas (areas other than the quality defect area selected by the user) in the target data unit is identified, and the quality defect area in other areas is obtained. The quality defect area selected by the user and the quality defect area in other areas are taken as the quality defect area in which the target data unit exists. The quality defect area in other areas has the same type of quality defect as the quality defect area selected by the user.
[0129] The matching feature is used to represent the relatively obvious and unchangeable attribute or quality of a quality defect type, and can be represented in the form of a feature vector. For example, whether a quality defect has the matching feature of a certain type of quality defect is determined by identification. If the quality defect has the matching feature associated with a certain type of quality defect, the quality defect type of the quality defect can be determined.
[0130] One quality defect type can be configured with at least one matching feature, and the association between the quality defect type and the matching feature can be configured according to actual needs, and the embodiment of the present application is not limited thereto.
[0131] Optionally, the computing device can utilize a template matching algorithm to calculate the similarity of the matching features between the quality defect region indicated by the first user instruction and each of the other regions, and identify, as the quality defect region in the other regions, a region in the other regions that has a similarity of the matching features greater than or equal to a preset threshold, based on the similarity of the matching features between the quality defect region indicated by the first user instruction and each of the other regions.
[0132] For example, the user can select the quality defect region to be optimized by selecting the quality defect region in the human-computer interaction interface, and S601 can be specifically implemented as: receiving a first user instruction, the first user instruction being used to indicate one or more quality defect regions; identifying, based on a matching feature corresponding to at least one quality defect type present in the quality defect region indicated by the first user instruction, other regions in the target data unit other than the quality defect region indicated by the first user instruction, to obtain a quality defect region in the other regions that has a quality defect indicated by the at least one quality defect type; and determining the quality defect region indicated by the first user instruction and the quality defect region in the other regions as at least one quality defect region present in the target data unit.
[0133] In implementation 2, the computing device automatically identifies the quality defect region present in the target data unit by machine.
[0134] In implementation 2, the computing device automatically identifies the quality defect region present in the target data unit by machine.
[0135] It should be understood that in implementation 2, the target data unit is a data unit at a preset position in the multimedia data.
[0136] The first identification strategy is used to identify a region in the multimedia data that has a quality defect, and the specific implementation of the first identification strategy can be configured according to actual needs, which is not limited in the embodiments of the present application.
[0137] For example, the first identification strategy can be an AI model. The first identification strategy is trained based on a large amount of sample data, and the training process is not limited in the embodiments of the present application.
[0138] Optionally, the first identification strategy is an AI model. Further, the first identification strategy can be a neural network model such as a multilayer perceptron model, a recurrent neural network model, a convolutional neural network model, a transformer model, etc.
[0139] For example, the first identification strategy is a target detection algorithm based on deep learning, such as a Faster R-CNN algorithm, a YOLO algorithm, etc.
[0140] In implementation 3, the user selects the data unit expected to optimize the quality defect, and the data unit operated by the user in the human-computer interaction interface is taken as the target data unit. Then, the computing device identifies the target data unit by using a first identification strategy in a manner of machine automatic identification, to obtain a quality defect area existing in the target data unit.
[0141] S602, the computing device determines at least one second optimization strategy based on the at least one quality defect area.
[0142] The second optimization strategy is used to optimize the quality defect in the at least one quality defect area existing in the target data unit.
[0143] Specifically, one second optimization strategy can be used to optimize one type of quality defect associated with itself, and different second optimization strategies can be pre-configured in association with the types of quality defects. In S602, the computing device can determine at least one second optimization strategy associated with at least one type of quality defect existing in the at least one quality defect area. That is, according to the type of quality defect in the at least one quality defect area existing in the target data unit, the second optimization strategy associated with the type of quality defect can be determined.
[0144] For example, there are N types of quality defects in the at least one quality defect area existing in the target data unit, and N second optimization strategies (one-to-one associated with the N types of quality defects) are determined in S602 to optimize the N types of quality defects in the at least one quality defect area existing in the target data unit.
[0145] For example, the second optimization strategy includes any one of the following: an AI model, a non-AI model for low-frequency filtering, a non-AI model for edge detection, or a non-AI model for color correction, or a non-AI model for optimizing other quality defects. The implementation form of the second optimization strategy is not limited in the embodiments of the present application.
[0146] Further, in S602, the computing device determines the type of quality defect in the quality defect area first, and then determines the second optimization strategy associated with the type of quality defect according to the quality defect area determined in S601.
[0147] Optionally, the computing device can determine the type of quality defect in the quality defect area in the following two ways, but is not limited to the following two ways:
[0148] Method one: the computing device automatically identifies the type of quality defect.
[0149] The computing device can employ a second identification strategy to identify the at least one quality defect region, to obtain at least one quality defect type existing in the at least one quality defect region.
[0150] The second identification strategy is configured to identify the quality defect type existing in the quality defect region. The specific implementation of the second identification strategy can be configured according to actual needs, and the embodiments of the present application are not limited thereto.
[0151] For example, the second identification strategy can be an AI model. The second identification strategy is trained based on a large amount of sample data, and the training process is not limited by the embodiments of the present application.
[0152] For example, the second identification strategy can be a multilayer perceptron (MLP) model, a recurrent neural network (RNN) model, a convolutional neural network (CNN) model, or a transformer model.
[0153] In the second way, the user specifies the quality defect type.
[0154] In the second way, the method provided by the present application further receives a second user instruction, the second user instruction being used to indicate the quality defect type; and determines the quality defect type indicated by the second user instruction as the at least one quality defect type existing in the at least one quality defect region.
[0155] The second user instruction can be input through a dialog box, selected in a menu, or other ways, and transmitted to the computing device. The embodiments of the present application are not limited to the input method of the second user instruction.
[0156] The above two ways of obtaining the quality defect type of the quality defect region are only examples and do not limit the process of obtaining the quality defect type of the quality defect region in the present application.
[0157] S603, the computing device optimizes the target data unit based on at least one second optimization strategy.
[0158] Specifically, the computing device optimizes the quality defect in the at least one quality defect region existing in the target data unit according to the at least one second optimization strategy.
[0159] Exemplarily, the second optimization strategy is used for the AI model, and a quality defect region in which the target data unit exists in the multimedia data can be cropped out and input into a second optimization strategy corresponding to a type of quality defect existing in the quality defect region. An output of the second optimization strategy is a quality defect region optimized by the second optimization strategy.
[0160] Exemplarily, when multiple types of quality defects exist in a quality defect region, data of the quality defect region in the multimedia data is input into a second optimization strategy corresponding to each type of quality defect existing in the quality defect region to obtain multiple optimized quality defect regions, and then data fusion is performed to obtain a final optimized quality defect region. The process of using the second optimization strategy to optimize different quality defect regions is the same, and will not be repeated here.
[0161] The scheme provided in the application is used to optimize the target data unit by using the second optimization strategy after the multimedia data is processed by the first optimization strategy, and the quality defects in the quality defect region in the target data unit are eliminated in a targeted manner. The second optimization strategy for solving the quality defects can be configured according to actual needs to ensure the accuracy of eliminating the quality defects by using the second optimization strategy. In this way, even if the multimedia data processed by the first optimization strategy introduces quality defects, the second optimization strategy can be used to process only the quality defect region (part of the region in the multimedia data) in which the quality defects exist in a targeted manner to eliminate the quality defects in the quality defect region. In addition, the second optimization strategy is used to process the quality defect region in a targeted manner, and the probability of introducing new quality defects is very low.
[0162] Further, after the target data unit is optimized by using the second optimization strategy in S603, the multimedia data in which the quality defects are eliminated can be obtained, as shown in FIG. 7, the method for processing the multimedia data provided in the embodiment of the application further includes the process of S604.
[0163] S604, the computing device fuses the target data unit optimized by the second optimization strategy into the multimedia data according to a position of the target data unit in the multimedia data.
[0164] Specifically, in S604, the target data unit in the multimedia data fused is the target data unit optimized by using the second optimization strategy in S603.
[0165] In a possible implementation, the computing device takes out the target data unit from the multimedia data in S603, and then optimizes the target data unit by using the second optimization strategy. In S604, the data unit at the position of the target data unit in the multimedia data is replaced by the target data unit optimized in S603.
[0166] For example, the multimedia data is a video stream, and the target data unit is the 5th image. The 5th image in the multimedia data is replaced by the optimized 5th image.
[0167] In another possible implementation, the computing device extracts the quality defect region in the target data unit from the multimedia data in S603, and obtains the optimized quality defect region by using the second optimization strategy. In S604, the optimized quality defect region obtained in S603 is filled into the position of the quality defect region in the target data unit in the multimedia data according to the size and proportion of the quality defect region before processing.
[0168] Of course, the fusion process in S604 can be configured according to actual needs, and depends on the processing manner of the target data unit in S603 using the second optimization strategy. As long as the target data unit in the multimedia data obtained by fusion in S604 is the target data unit optimized by using the second optimization strategy in S603, it belongs to the purpose of fusion described in S604.
[0169] For example, the system architecture of the scheme provided in the present application can be as shown in FIG. 8. The system architecture can include a basic optimization module, a user interaction module, a defect identification module, a defect matching module, a defect processing module, and a defect processing model library. The function descriptions of the modules in the system architecture shown in FIG. 8 are as follows:
[0170] The defect processing model library records the second optimization strategy associated with the quality defect type and the matching feature associated with the quality defect type.
[0171] The basic optimization module is used to optimize the multimedia data by using the first optimization strategy.
[0172] The user interaction module is used to receive user operations, and determine the quality defect region to be optimized according to the interactive information.
[0173] The defect identification module is used to identify the quality defect type from the quality defect region.
[0174] The defect matching module is used to match the quality defect region in other regions according to the matching feature associated with the quality defect type by using the matching manner.
[0175] The defect processing module is used to obtain the second optimization strategy associated with the quality defect type from the defect processing model library, and optimize the quality defect region determined by the user interaction module and the quality defect region determined by the defect matching module by using the second optimization strategy, and output the multimedia data in which the defects are eliminated.
[0176] Exemplarily, the system architecture diagram of the scheme provided in the present application can be as shown in FIG. 9. The system architecture can include a basic optimization module, a defect automatic identification module, a defect processing module, and a defect processing model library. In the system architecture shown in FIG. 9, the process of user interaction shown in FIG. 8 is replaced by defect automatic identification. The function descriptions of the various modules in the system architecture shown in FIG. 9 are as follows:
[0177] The defect processing model library records the second optimization strategy associated with the quality defect type.
[0178] The basic optimization module is used to optimize the multimedia data by using the first optimization strategy.
[0179] The defect automatic identification module is trained by using the defect data set, and is used to automatically identify the quality defect area and the quality defect type from the target data unit.
[0180] The defect processing module is used to obtain the second optimization strategy associated with the quality defect type from the defect processing model library, and to optimize the quality defect area determined by the defect automatic identification module by using the second optimization strategy, and to output the multimedia data in which the defects are eliminated.
[0181] Further, the description in the foregoing S601 describes that the quality defect area in which the quality defect exists in the target data unit is determined through user interaction, or the implementation of the target data unit is determined. The process of user interaction is exemplarily described below through specific examples.
[0182] Exemplarily, it is assumed that the multimedia data is a video stream, and the video stream to be edited is stored in the photo album APP of the terminal used by the user. The video stream to be edited has undergone noise reduction AI processing (processing of the first optimization strategy).
[0183] The user operates the terminal to further optimize the quality of the video stream to be edited by using the scheme provided in the present application. As shown in (a) of FIG. 10, the user clicks the icon 1001 of the video optimization APP to enter the interface of the video optimization APP as shown in (b) of FIG. 10. The interface shown in (b) of FIG. 10 includes an adding icon 1002. The user clicks the icon 1002 to enter the interface shown in (c) of FIG. 10, and the interface presents the icons (which can be in the form of thumbnails) of all multimedia data in the photo album of the terminal, including the icon 1003 of the video stream to be edited.
[0184] Then, the user clicks the icon 1003 of the video stream to be edited in (c) of FIG. 10 to further optimize the quality of the video stream to be edited by using the scheme provided in the present application.
[0185] Example 1, the user clicks the icon 1003 of the video stream to be edited in (c) of FIG. 10, enters the video editing interface shown in (d) of FIG. 10, and the interface presents each image frame (data unit) of the video stream to be edited. The user can slide to view each image frame, and selects the quality defect area to be optimized. For example, the user selects the area 1 as the quality defect area by the box selection operation (the first user instruction) shown in the interface in (d) of FIG. 10. The user can select the quality defect area to be optimized in each image frame to be optimized by the same operation.
[0186] After example 1, the case one or case two described in the implementation one of S601 can be used to determine the quality defect area with quality defects in the target data unit, and then the processes of S602 and S603 are performed to optimize the target data unit.
[0187] Example 2, the user clicks the icon 1003 of the video stream to be edited in (c) of FIG. 10, enters the video editing interface shown in (a) of FIG. 11, and the interface presents each image frame (data unit) of the video stream to be edited. The user can slide to view each image frame, and selects the quality defect area to be optimized. For example, the user selects the area 1 as the quality defect area by the box selection operation (the first user instruction) shown in the interface in (a) of FIG. 11. Further, the user can enter the type selection interface shown in (b) of FIG. 11 by the long press operation in the area 1, and the interface presents the alternative quality defect types. In the type selection interface shown in (b) of FIG. 11, the user selects one or more quality defect types by the click operation as the quality defect types in the quality defect area of the area 1. The user can select the quality defect area and the quality defect types to be optimized in each image frame to be optimized by the same operation.
[0188] In example 2, the long press operation and the click operation can be understood as the second user instruction.
[0189] After example 2, the case one described in the implementation one of S601 can be used to determine the quality defect area with quality defects in the target data unit, and then the processes of S602 and S603 are performed to optimize the target data unit.
[0190] In example 3, the user clicks the icon 1003 of the video stream to be edited in (c) of FIG. 10, and enters the video editing interface shown in FIG. 12, in which the image frames (data units) of the video stream to be edited are presented. The user can slide to view the image frames, and select the data units to be optimized as the target data units to be optimized. For example, in the interface shown in FIG. 12, a selection box is configured in each image frame, and the image frames selected by the user are the first, second and fourth image frames. Further, the user can long press any image frame, and a selection box appears in the lower right corner of the image frame in the interface.
[0191] After example 3, the description of implementation 3 in S601 above can be used to determine the quality defect area with quality defects in the target data unit, and then the processes of S602 and S603 are performed to optimize the target data unit.
[0192] It can be understood that the apparatus for processing multimedia data in the embodiments of the present application can perform some or all of the steps in the embodiments of the present application, and these steps or operations are only examples, and the embodiments of the present application can also perform other operations or various modifications of the operations. In addition, each step can be performed in a different order from the order presented in the embodiments of the present application, and it is possible that not all operations in the embodiments of the present application are performed.
[0193] In various embodiments of the present application, the terms and / or descriptions of different embodiments are consistent and can be mutually referred to if there is no special description and no logical conflict, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.
[0194] It can be understood that the methods and / or steps realized by the apparatus for processing multimedia data in each of the above embodiments can also be realized by components (such as chips or circuits) that can be used in the apparatus.
[0195] Alternatively, it can be understood that the apparatus for processing multimedia data contains hardware structures and / or software modules corresponding to each function in order to realize the above functions. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed herein, the present application can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is realized in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to realize the described functions, but such implementation should not be considered beyond the scope of the present application.
[0196] The embodiments of the present application can divide the device for processing multimedia data in the above-mentioned method embodiments into functional modules. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The integrated module can be implemented in the form of hardware or software functional module. It should be noted that the division of the modules in the embodiments of the present application is illustrative, and is only a logical function division. In actual implementation, another division manner can be used.
[0197] For example, FIG. 13 shows a structural schematic diagram of a device 130 for processing multimedia data. The device 130 for processing multimedia data includes a first determining module 1301, a second determining module 1302 and an optimization module 1303.
[0198] The first determining module 1301 is configured to perform S601 in the processes of FIG. 6 and FIG. 7, the second determining module 1302 is configured to perform S602 in the processes of FIG. 6 and FIG. 7, and the optimization module 1303 is configured to perform S603 in the processes of FIG. 6 and FIG. 7.
[0199] Optionally, the device 130 for processing multimedia data of the present application can further include a fusion module 1304 configured to perform S604 in the process of FIG. 7.
[0200] Optionally, one second optimization strategy is used to optimize a type of quality defects associated therewith, and the second determining module is specifically configured to determine at least one second optimization strategy associated with at least one type of quality defects existing in at least one quality defect region.
[0201] Optionally, the first determining module is specifically configured to: receive a first user instruction used to indicate one or more quality defect regions; and determine the quality defect region indicated by the first user instruction as at least one quality defect region existing in the target data unit.
[0202] Optionally, different quality defect types correspond to different matching features, and the first determining module is specifically configured to: receive a first user instruction used to indicate one or more quality defect regions; identify other regions in the target data unit except the quality defect region indicated by the first user instruction based on matching features corresponding to at least one type of quality defects existing in the quality defect region indicated by the first user instruction, to obtain a quality defect region in which at least one type of quality defects indicated by the first user instruction exists; and determine the quality defect region indicated by the first user instruction and the quality defect region in the other regions as at least one quality defect region existing in the target data unit.
[0203] Optionally, the first determining module is specifically configured to: adopt a first identification strategy to identify the target data unit to obtain the at least one quality defect area in the target data unit.
[0204] Optionally, the device for processing multimedia data can further include an identification module configured to adopt a second identification strategy to identify the at least one quality defect area to obtain the at least one quality defect type existing in the at least one quality defect area.
[0205] Optionally, the device for processing multimedia data can further include a second receiving unit and a second determining unit. The second receiving unit is configured to receive a second user instruction for indicating a quality defect type. The second determining unit is configured to determine the quality defect type indicated by the second user instruction as the at least one quality defect type existing in the at least one quality defect area.
[0206] In a possible implementation, the second optimization strategy includes any one of the following: an AI model, a non-AI model for low-frequency filtering, a non-AI model for edge detection, or a non-AI model for color correction.
[0207] In another possible implementation, the quality defect type includes any one of the following: loss of detail smearing, inter-frame inconsistency of texture representation, noise flicker, bright noise, color noise, color edge, contour convexity defect, color oversaturation, ghosting, pseudo color, stroboscopic, or color cast.
[0208] All the related contents of the steps involved in the method embodiments described above can be cited from the function description of the corresponding function modules, and will not be repeated here.
[0209] FIG. 14 is a schematic diagram of a hardware structure of a server according to an embodiment of the present application. As shown in FIG. 14, the server 1400 can include a processor 1401, a memory 1402, and a communication interface 1403; and the various components are coupled together through a communication line 1404. It can be understood that the communication line 1404 is used to realize the connection and communication between the components. In addition to including a data bus, the communication line 1404 also includes a power supply bus, a control bus, and a state signal bus. However, for the purpose of clear illustration, all the buses are marked as the communication line 1404 in FIG. 14.
[0210] The processor 1401 is configured to execute the steps of the method in any of the preceding embodiments when running the computer program.
[0211] The memory 1402 is configured to store the computer program capable of running on the processor 1401.
[0212] The communication interface 1403 is configured to receive and send signals in the process of transceiving information with other external network elements.
[0213] It is to be appreciated that the memory 1402 in embodiments of the application can be volatile, nonvolatile, or a combination of both. The non-volatile memory can be, for example, read only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically EPROM (EEPROM), or flash memory. The volatile memory can be, for example, random access memory (RAM), which acts as external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), SynchBurst SDRAM (SLDRAM), and direct rambus RAM (DRRAM). The memory 1402 of the subject systems and methods is intended to include, without being limited to, these and any other suitable types of memory.
[0214] The processor 1401 can be an integrated circuit chip having a processing capability for signals. In implementation, each step of the above method can be completed by integrated logic circuits or instructions in software form in the processor 1401. The processor 1401 described above can be a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. Each method, step and logic block diagram disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register, and other mature storage media in the art. The storage medium is located in the storage 1402, and the processor 1401 reads the information in the storage 1402 and combines the hardware to complete the steps of the above method.
[0215] Optionally, as another embodiment, the processor 1401 is further configured to execute the steps of the method of any one of the preceding embodiments when the computer program is run.
[0216] In a specific implementation, as an embodiment, the processor 1401 can include one or more CPUs, such as CPU0 and CPU1 in FIG. 14.
[0217] It should be noted that FIG. 14 is only an example of a server and does not limit the specific structure of the server.
[0218] As another form of the embodiment, a computer readable storage medium is provided, and the computer readable storage medium stores computer instructions or programs. When the computer instructions or programs are run on a computer, the method for processing multimedia data in any one of the method embodiments is executed.
[0219] As another form of the embodiment, a computer program product is provided, and the computer program product includes computer programs or instructions. When the computer programs or instructions are executed by a processor, the method for processing multimedia data in any one of the method embodiments is executed.
[0220] The embodiment of the present application further provides a chip comprising a processor, wherein the processor is configured to execute computer instructions to perform any of the above method embodiments. In a possible design, the chip further comprises a memory configured to store the computer instructions.
[0221] The method steps in the embodiments can be implemented by hardware, or by a combination of software and hardware. The software instructions can be composed of one or more pieces of software modules. The software modules can be stored in any suitable storage medium or device, such as a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically EPROM (EEPROM), a register, a hard disk, a floppy disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, such that the processor can read information from, and write information to, the storage medium. Of course, the storage medium can be a component of the processor. The processor and the storage medium can be located in an ASIC. The ASIC can be located in a network device or a terminal device. Of course, the processor and the storage medium can also exist as a separate component in a network device or a terminal device.
[0222] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks, SSDs).
[0223] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, the disclosure, and the appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.
[0224] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.
[0225] In the description of the present application, unless otherwise specified, " / " means that the objects before and after the correlation are in an "or" relationship, for example, A / B can represent A or B; "and / or" in the present application is only a description of the correlation of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. In the description of the present application, unless otherwise specified, "multiple" means two or more than two. "At least one of the following (one) or the like means any combination of the items, including any combination of single (one) or multiple items. For example, at least one of a, b and (or) c can represent: a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0226] In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, "first", "second" and the like are used to distinguish the same items or similar items with basically the same function and role. The skilled in the art can understand that "first", "second" and the like do not limit the quantity and execution order, and "first", "second" and the like do not necessarily mean different.
[0227] In the embodiments of the present application, the words "exemplarily" or "for example" are used to represent as an example, illustration or explanation. Any embodiment or design scheme described as "exemplarily" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words "exemplarily" or "for example" are intended to present the relevant concept in a specific manner for understanding.
[0228] It can be understood that the "embodiments" mentioned throughout the specification mean that the specific features, structures or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, throughout the specification, various embodiments do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It can be understood that in various embodiments of the present application, the size of the sequence number of each process does not mean the execution order, and the execution order of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0229] Some optional features in the embodiments of the present application can be implemented independently in some scenarios without relying on other features, such as the scheme currently based on, to solve the corresponding technical problems and achieve the corresponding effects, and can also be combined with other features according to needs in some scenarios. Correspondingly, the devices given in the embodiments of the present application can also correspondingly implement these features or functions, which will not be described here.
[0230] In the present application, the same or similar parts between various embodiments can be mutually referred to, unless otherwise specified. In various embodiments of the present application, the terms and / or descriptions between different embodiments are consistent and can be mutually referred to, unless otherwise specified and logically conflicted. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship. The following embodiments of the present application do not constitute a limitation on the protection scope of the present application.
Claims
1. A method of processing multimedia data, characterized by, The method comprises: determining at least one quality defect region existing in a target data unit in multimedia data processed by using a first optimization strategy, the target data unit being one or more data units in the multimedia data; the multimedia data comprising any one of the following types: image, video, or audio; determining at least one second optimization strategy based on the at least one quality defect region; optimizing the target data unit based on the at least one second optimization strategy.
2. The method of claim 1, wherein, One second optimization strategy is used to optimize a type of quality defect associated therewith, and the determining of the at least one second optimization strategy based on the at least one quality defect region comprises: determining the at least one second optimization strategy associated with at least one quality defect type existing in the at least one quality defect region.
3. The method according to claim 1 or 2, characterized in that, The determining of the at least one quality defect region existing in the target data unit in the multimedia data processed by using the first optimization strategy comprises: receiving a first user instruction, the first user instruction being used to indicate one or more quality defect regions; determining the quality defect region indicated by the first user instruction as the at least one quality defect region existing in the target data unit.
4. The method according to claim 1 or 2, characterized in that, Different quality defect types correspond to different matching features, and the determining of the at least one quality defect region existing in the target data unit in the multimedia data processed by using the first optimization strategy comprises: receiving a first user instruction, the first user instruction being used to indicate one or more quality defect regions; identifying other regions in the target data unit other than the quality defect region indicated by the first user instruction based on matching features corresponding to at least one quality defect type existing in the quality defect region indicated by the first user instruction, to obtain a quality defect region in the other regions in which quality defects indicated by the at least one quality defect type exist; determining the quality defect region indicated by the first user instruction and the quality defect region in the other regions as the at least one quality defect region existing in the target data unit.
5. The method of claim 1, wherein the determining of the at least one quality defect region existing in the target data unit in the multimedia data processed by using the first optimization strategy comprises: identifying the target data unit by using a first identification strategy to obtain the at least one quality defect region.
6. The method according to claim 4 or 5, characterized in that, The method further comprises: identifying at least one quality defect type existing in the at least one quality defect region by using a second identification strategy.
7. The method according to claim 3 or 4, characterized in that, The method further comprises: receiving a second user instruction, the second user instruction being used to indicate a quality defect type; determining the quality defect type indicated by the second user instruction as the at least one quality defect type existing in the at least one quality defect region.
8. The method according to any one of claims 1 to 7, characterized in that, The method further comprises: fusing the target data unit optimized by using the second optimization strategy to the multimedia data according to a position of the target data unit in the multimedia data.
9. The method according to any one of claims 1 to 8, characterized in that, The second optimization strategy comprises any one of the following: an AI model, a non-AI model for low-frequency filtering, a non-AI model for edge detection, or a non-AI model for color correction.
10. The method of claim 2, wherein, The quality defect type includes any one of the following: detail painting loss, texture performance inter-frame inconsistency, noise flicker, bright noise, color noise, color edge, contour convexity defect, color oversaturation, ghosting, pseudo color, stroboscopic, or color cast.
11. An apparatus for processing multimedia data, the apparatus comprising: The device includes: A first determining module is configured to determine at least one quality defect area existing in a target data unit after the multimedia data is processed by using a first optimization strategy, the target data unit being one or more data units in the multimedia data, and the multimedia data including any one of the following types: image, video, or audio. A second determining module is configured to determine at least one second optimization strategy based on the at least one quality defect area. An optimization module is configured to optimize the target data unit based on the at least one second optimization strategy.
12. The apparatus of claim 11, wherein, One second optimization strategy is used to optimize a type of quality defect associated therewith, and the second determining module is specifically configured to determine the at least one second optimization strategy associated with at least one quality defect type existing in the at least one quality defect area.
13. The apparatus of claim 11 or 12, wherein, The first determining module is specifically configured to: receive a first user instruction, the first user instruction being used to indicate one or more quality defect areas; and determine the quality defect areas indicated by the first user instruction as the at least one quality defect area existing in the target data unit.
14. The apparatus of claim 11 or 12, wherein, Different quality defect types correspond to different matching features, and the first determining module is specifically configured to: receive a first user instruction, the first user instruction being used to indicate one or more quality defect areas; identify other areas in the target data unit except for the quality defect areas indicated by the first user instruction based on matching features corresponding to at least one quality defect type existing in the quality defect areas indicated by the first user instruction, to obtain quality defect areas in the other areas in which quality defects indicated by the at least one quality defect type exist; and determine the quality defect areas indicated by the first user instruction and the quality defect areas in the other areas as the at least one quality defect area existing in the target data unit. The first determining module is specifically configured to: identify the target data unit by using a first identification strategy to obtain the at least one quality defect area.
15. The apparatus of claim 11, wherein, The device further includes:
16. The apparatus of claim 14 or 15, wherein, An identification module is configured to identify the at least one quality defect area by using a second identification strategy to obtain at least one quality defect type existing in the at least one quality defect area. The device further includes:
17. The apparatus of claim 13 or 14, wherein, A second receiving unit is configured to receive a second user instruction, the second user instruction being used to indicate a quality defect type. A second determining unit is configured to determine the quality defect type indicated by the second user instruction as at least one quality defect type existing in the at least one quality defect area. The device further includes:
18. The apparatus of any of claims 11-17, wherein, a fusion module configured to fuse the target data unit optimized by the second optimization strategy into the multimedia data according to a position of the target data unit in the multimedia data.
19. The apparatus of any of claims 11-18, wherein, The second optimization strategy comprises any one of an AI model, a non-AI model for low-frequency filtering, a non-AI model for edge detection, or a non-AI model for color correction.
20. The apparatus of claim 12, wherein, The quality defect type comprises any one of detail smearing loss, texture representation inter-frame inconsistency, noise flicker, bright noise, color noise, color edge, contour convexity defect, color oversaturation, ghosting, pseudo color, stroboscopic, or color cast.
21. A server, comprising: The server comprises a processor and a memory, the memory being configured to store computer instructions, when the processor executes the computer instructions, the server is caused to execute the method of any one of claims 1-10.
22. An electronic device, comprising: The electronic device comprises a processor and a memory, the memory being configured to store computer instructions, when the processor executes the computer instructions, the electronic device is caused to execute the method of any one of claims 1-10.
23. A chip, characterized by The chip comprises a processor configured to execute computer instructions to execute the method of any one of claims 1-10.
24. The chip of claim 23, wherein, The chip further comprises a memory configured to store the computer instructions.
25. A computer program product, characterised in that, The computer program product comprises computer programs or instructions, when the computer programs or instructions are executed by a processor, the steps of the method of any one of claims 1-10 are implemented.
26. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions or programs, when the computer instructions or programs are run on a computer, the method of any one of claims 1-10 is implemented.
Citation Information
Patent Citations
Image reconstruction system and method
CN115605915A
Image processing method and device, computer readable storage medium and terminal
CN117474792A
Image quality optimization method and device, electronic equipment, medium and program product
CN118037602A
Audio data processing method and apparatus, device, and storage medium
US20230146871A1