Coding and decoding method and related device

By providing a unified encoding and decoding capability interface in the operating system, applications at the application layer can directly call atomized encoding and decoding capabilities, solving the problem of low encoding and decoding efficiency of electronic devices and improving user experience.

CN120295693APending Publication Date: 2025-07-11HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410744805.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2024-06-07
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the low encoding and decoding efficiency of electronic devices leads to slow data storage speed, stuttering video playback, and out-synchronization of audio and video, which affects the user experience.

Method used

It provides a unified encoding and decoding capability interface, encapsulating atomized encoding and decoding capabilities through the media data interface of the operating system. Applications at the application layer can be called directly, reducing the attention to encoding and decoding implementation.

Benefits of technology

It improves the encoding and decoding efficiency, reduces the complexity of application development, improves the user experience, and solves the problem of low encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295693A_ABST
    Figure CN120295693A_ABST
Patent Text Reader

Abstract

The invention discloses a coding and decoding method and a related device. The electronic equipment can provide simple functional interfaces for each application of the application layer, and the application can directly call the functional interfaces to obtain the atomization coding and decoding capability provided by an operating system of the electronic equipment. According to the scheme, the operating system undertakes more encoding and decoding work, the application does not need to pay attention to the implementation of encoding and decoding, and the method is more friendly to the application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of terminals, and in particular to encoding and decoding methods and related devices. Background Art

[0002] Electronic devices such as mobile phones, tablet computers, and laptop computers can generate and store data, and the types and formats of this data are increasing. Users can obtain various information through the data in the electronic device. The encoding and decoding efficiency of the data affects the user experience. For example, low encoding efficiency affects the storage speed of the data, and low decoding efficiency causes problems such as video playback stuttering, audio-video out-of-sync, and low frame rate during video playback. Improving the encoding and decoding speed of the data is an important direction for improving the user experience. Summary of the Invention

[0003] This application provides encoding and decoding methods and related devices. The electronic device can provide unified encoding and decoding capabilities for applications to use, without the applications having to concern themselves with the implementation of encoding and decoding.

[0004] In a first aspect, an encoding method is provided. This method is applied to a first electronic device on which a first operating system is running. The method may include: running a first application on the first operating system, where the first application is installed on the first electronic device, and the first operating system provides: a first interface, a media data interface, and a second capability; the first application is a third-party application, and the first application includes program code for calling the first interface, where the first interface is used to call the media data interface, and the media data interface encapsulates the second capability. The second capability includes: the ability to perform encoding on audio data, the ability to perform encoding on video data, and the ability to perform encoding on audio-video data. Among them, the ability to perform encoding on audio data includes multiple atomic encoding capabilities for audio, the ability to perform encoding on video data includes multiple atomic encoding capabilities for video, and the ability to perform encoding on audio-video data includes multiple atomic encoding capabilities for audio-video; detecting a user operation to start a first function of the first application, where the first function is used to trigger the encoding of first media data; encoding the first media data.

[0005] Through the method of the first aspect, the electronic device can provide a simple first interface to each application in the application layer. The application can directly call this first interface to obtain the atomic encoding capabilities provided by the operating system of the electronic device. In such a solution, the operating system undertakes more encoding and decoding work, and there is no need for the application to concern itself with the implementation of encoding, which is more user-friendly for the application.

[0006] In combination with the first aspect, in some embodiments, when the first electronic device encodes the first media data, the first application may call the first interface, and then the first interface calls the media data interface, and the media data interface calls one or more atomized encoding capabilities in the second capabilities to implement the encoding of the first media data.

[0007] In combination with the first aspect, in some embodiments, the capabilities for performing encoding on audio data include: a first client capability and a first server capability. The first client capability is used to transmit the indication information and parameter content of the parameters for performing audio encoding passed in by the first application to the first server capability, and the first server capability is used to select one or more atomized capabilities from multiple atomized audio encoding capabilities to encode the first media data according to the information passed in by the first application; the capabilities for performing encoding on video data include: a second client capability and a first server capability. The second client capability is used to transmit the indication information and parameter content of the parameters for performing audio encoding passed in by the first application to the first server capability, and the first server capability is used to select one or more atomized capabilities from multiple atomized video encoding capabilities to encode the first media data according to the information passed in by the first application; the capabilities for performing encoding on audio-visual data include: a third client capability and a first server capability. The third client capability is used to transmit the indication information and parameter content of the parameters for performing audio encoding passed in by the first application to the first server capability, and the first server capability is used to select one or more atomized capabilities from multiple atomized audio-visual encoding capabilities to encode the first media data according to the information passed in by the first application. In this way, the electronic device can provide encoding capabilities in a client-server manner.

[0008] In some embodiments, the first server capability is further used to return the encoding result of the first media data to the first client capability or the second client capability or the third client capability, and the first client capability or the second client capability or the third client capability is further used to return the encoding result of the first media data to the first application.

[0009] In combination with the first aspect, in some embodiments, the first electronic device may further decode the first media data. Specifically, the method may further include: running a second application on the first operating system, the second application being installed on the first operating system, and the first operating system further providing: a second interface, the second application being a third-party application, the second application including program code for invoking the second interface, the second interface being used to start a fourth function provided by the second application, the second interface being used to invoke a media data interface, the media data interface encapsulating a second capability, the second capability further including: the capability to perform decoding on audio data, the capability to perform decoding on video data, the capability to perform decoding on audio-video data, wherein the capability to perform decoding on audio data includes multiple atomic decoding capabilities for audio decoding, the capability to perform decoding on video data includes multiple atomic decoding capabilities for video decoding, and the capability to perform decoding on audio-video data includes multiple atomic decoding capabilities for audio-video decoding; detecting a user operation to start the fourth function of the first application, the fourth function being used to trigger the decoding of the first media data; and decoding the first media data.

[0010] In this way, the electronic device can provide a simple second interface to each application in the application layer. The application can directly invoke this second interface to obtain the atomic decoding capabilities provided by the operating system of the electronic device. In such a solution, the operating system undertakes more encoding and decoding work, and there is no need for the application to concern itself with the implementation of decoding, which is more friendly to the application.

[0011] In combination with the first aspect, in some embodiments, when the first electronic device decodes the first media data, the first application may invoke the second interface, and then the second interface may invoke the media data interface, and the media data interface may invoke one or more atomic decoding capabilities in the second capability to implement the decoding of the first media data.

[0012] In combination with the first aspect, in some embodiments, the ability to perform decoding on audio data includes: a first client ability and a first server ability. The first client ability is used to transmit to the first server ability an indication information and parameter content of parameters for performing audio decoding passed in by a first application. The first server ability is used to select one or more atomic abilities for audio decoding from multiple atomic abilities for audio decoding according to the information passed in by the first application to decode the first media data. The ability to perform decoding on video data includes: a second client ability and a first server ability. The second client ability is used to transmit to the first server ability an indication information and parameter content of parameters for performing audio decoding passed in by the first application. The first server ability is used to select one or more atomic abilities for video decoding from multiple atomic abilities for video decoding according to the information passed in by the first application to decode the first media data. The ability to perform decoding on audio-visual data includes: a third client ability and a first server ability. The third client ability is used to transmit to the first server ability an indication information and parameter content of parameters for performing audio decoding passed in by the first application. The first server ability is used to select one or more atomic abilities for audio-visual decoding from multiple atomic abilities for audio-visual decoding according to the information passed in by the first application to decode the first media data. In this way, the electronic device can provide the decoding ability in a client-server manner.

[0013] In combination with the first aspect, in some embodiments, the first server ability is further used to return the result of decoding the first media data to the first client ability or the second client ability or the third client ability, and the first client ability or the second client ability or the third client ability is further used to return the result of decoding the first media data to the first application.

[0014] In combination with the first aspect, in some embodiments, the first media data is decoded by a hardware codec or a software codec of the first electronic device.

[0015] In combination with the first aspect, in some embodiments, when the first media data is decoded by a hardware codec of the first electronic device, the decoupling relationship between the buffer and the view can be decoupled. For example, the method may further include: storing the decoded data in a first buffer; outputting the data in the first buffer to a first view displayed on the display screen of the first electronic device for completing the playback of the data; detecting that the first buffer is damaged; outputting the data in a second buffer to the first view for completing the playback of the data.

[0016] In some embodiments, if the duration of damage to the first buffer is less than a first value, the second buffer is the buffer in the buffers for storing decoded data that is most similar to the first buffer; if the duration of damage to the first buffer is greater than the first value, the second buffer is the buffer in the buffers for storing decoded data that best matches the first view. The first value can be preset and will not be specifically limited here.

[0017] In some embodiments, if the first media data is in a format not supported by the atomization ability in the second capability, before decoding the first media data, the first electronic device converts the first media data from the first format to the second format, and then the first media data in the second format can be decoded. The second format is a format supported by the atomization ability in the second capability.

[0018] Among them, the transcoding methods include two types: 1. Transcoding of the same type. For example, both the first format and the second format are formats of a first data type (such as text, picture, video, audio, or audio-video), and the first data type is the type of the first media data. 2. Upward classification transcoding. For example, the first format is a format of a first data type (such as video), and the second format is a format of a second data type (such as text), and the data volume of the data of the second data type is less than the data volume of the data of the first data type.

[0019] In some embodiments, the first electronic device can also adjust the bit rate of the screen projection content. Specifically, the first electronic device receives the screen projection content of the second electronic device and the screen projection content of the third electronic device, plays the screen projection content of the second electronic device at a first bit rate, and plays the screen projection content of the third electronic device at a second bit rate. The first bit rate is less than the second bit rate, and the sum of the first bit rate and the second bit rate is equal to the total bit rate of the first electronic device; among them, the screen projection content of the second electronic device is static content, and the screen projection content of the third electronic device is dynamic content; or, the screen projection content of the second electronic device is low-bit content, and the screen projection content of the third electronic device is dynamic content; or, the screen projection content of the second electronic device is low-bit content, and the screen projection content of the third electronic device is static content.

[0020] In some embodiments, if it is detected that the screen focus is on the screen projection content of the first electronic device, the screen projection content of the second electronic device is played at a third bit rate, and the screen projection content of the third electronic device is played at a fourth bit rate. The third bit rate is greater than the fourth bit rate, and the sum of the third bit rate and the fourth bit rate is equal to the total bit rate of the first electronic device.

[0021] In combination with the first aspect, in some embodiments, the first electronic device may also perform hierarchical encoding on the first media data. Specifically, the first electronic device may obtain the data stream of the first media data; divide the data stream into n segments, and further divide each segment of the data stream into k parts, where n is greater than or equal to 1 and k is greater than or equal to 2, and the length of the data stream is the duration required for encoding the data stream; perform parallel encoding on the n segments of the data stream. Among them, the encoding start times of the first part of the data streams in the n segments of the data stream are the same. In each segment of the data stream, the encoding start time of the i-th part of the data stream is earlier than that of the (i + 1)-th part of the data stream by timeset, and timeset is less than or equal to the length of each part of the data stream. This can save the overall encoding time and improve the decoding efficiency, so as to play media files more efficiently and with higher quality.

[0022] In a second aspect, there is provided an electronic device, including: a memory, a processor, and a computer program stored on the memory. The computer program includes a first operating system. The first electronic device is installed with a first application. The first operating system provides: a first interface, a media data interface, and a second capability; the first application is a third-party application, and the first application includes program code for calling the first interface. The first interface is used to call the media data interface, and the media data interface encapsulates the second capability. The second capability includes: the ability to perform encoding on audio data, the ability to perform encoding on video data, and the ability to perform encoding on audio-visual data. Among them, the ability to perform encoding on audio data includes multiple atomic capabilities of audio encoding, the ability to perform encoding on video data includes multiple atomic capabilities of video encoding, and the ability to perform encoding on audio-visual data includes multiple atomic capabilities of audio-visual encoding; the processor executes the computer program to implement the method provided in the first aspect or any one of the embodiments of the first aspect.

[0023] In a third aspect, there is provided a computer-readable storage medium, on which a computer program is stored. The computer program includes a first operating system. The first electronic device is installed with a first application. The first operating system provides: a first interface, a media data interface, and a second capability; the first application is a third-party application, and the first application includes program code for calling the first interface. The first interface is used to call the media data interface, and the media data interface encapsulates the second capability. The second capability includes: the ability to perform encoding on audio data, the ability to perform encoding on video data, and the ability to perform encoding on audio-visual data. Among them, the ability to perform encoding on audio data includes multiple atomic capabilities of audio encoding, the ability to perform encoding on video data includes multiple atomic capabilities of video encoding, and the ability to perform encoding on audio-visual data includes multiple atomic capabilities of audio-visual encoding; when the computer program is executed by the processor, it implements the method provided in the first aspect or any one of the embodiments of the first aspect.

[0024] Fourthly, a computer program product is provided. The computer program product includes a computer program, and the computer program includes a first operating system. A first electronic device is installed with a first application. The first operating system provides: a first interface, a media data interface, and a second capability. The first application is a third-party application, and the first application includes program code for calling the first interface. The first interface is used to call the media data interface, and the media data interface encapsulates the second capability. The second capability includes: the ability to encode audio data, the ability to encode video data, and the ability to encode audio-visual data. Among them, the ability to encode audio data includes multiple atomic capabilities of audio encoding, the ability to encode video data includes multiple atomic capabilities of video encoding, and the ability to encode audio-visual data includes multiple atomic capabilities of audio-visual encoding. When the computer program is executed by a processor, the method provided in the first aspect or any one of the implementation manners of the first aspect is implemented. Description of the Drawings

[0025] Figure 1 It is a schematic diagram of a media service framework;

[0026] Figure 2 It is a software architecture diagram of the electronic device provided in the embodiment of the present application;

[0027] Figure 3 It is a flowchart of hardware encoding and decoding provided in the embodiment of the present application;

[0028] Figure 4A Shows based on Figure 1 The encoding and decoding process of the media service framework shown;

[0029] Figure 4B Shows based on Figure 2 The encoding and decoding process of the software architecture shown;

[0030] Figure 5A Shows the corresponding relationship between the buffer and the view in the resource pool;

[0031] Figure 5B Shows the display method when the buffer is damaged or abnormal;

[0032] Figure 6 It is a flowchart of software encoding and decoding provided in the embodiment of the present application;

[0033] Figure 7 It is a transcoding flowchart provided in the embodiment of the present application;

[0034] Figure 8 It is a transcoding flowchart of a video in HDR vivid format;

[0035] Figure 9 Shows a screen mirroring scenario;

[0036] Figure 10 shows the processing flow of the audio - video stream provided by the embodiments of the present application;

[0037] Figure 11A shows the single - layer streaming media encoding method;

[0038] Figure 11B and Figure 11C shows the layered streaming media encoding method;

[0039] Figure 12 is the flowchart of the encoding and decoding method provided by the embodiments of the present application;

[0040] Figure 13 is the hardware structure block diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manner

[0041] Data can also be referred to as media data. Data can be divided into multiple categories, for example, it can include audio data, video data, audio - video data, picture data, text data, etc. Audio data can include audio data, video data includes multiple frames of image data, audio - video data includes both audio data and image data, picture data can include picture data, and text data can include text data.

[0042] The operation process of an electronic device often involves the encoding and decoding of data. An electronic device can encode audio data, image data, picture data, text data, etc. to generate data. The electronic device can also perform a decoding operation on the data to obtain the original data and then play the original data for the user to view.

[0043] Encoding is the process of converting data from one form or format to another form or format. The main purpose is to convert the original data into a digital signal in a specified format. The purpose of encoding is to reduce the data volume, simplify data transmission and storage, and provide a standardized data format for applications. Decoding is the reverse process of encoding and is used to restore the original data content.

[0044] The encoding and decoding efficiency of data affects the user experience. For example, low encoding efficiency affects the data storage speed, and low decoding efficiency causes problems such as video playback stuttering, audio - video out - of - sync, and low video playback frame rate. It can be seen that improving the encoding and decoding efficiency is a way to improve the user experience.

[0045] Figure 1 Exemplarily shows the structure of the media service framework provided by Android. As Figure 1As shown in the figure, Android's media service framework includes a framework layer, which includes the following three service components: media codec (MediaCodec), media data extractor (MediaExtractor) and media encapsulator (MediaMuxer). MediaCodec is responsible for encoding and decoding. MediaExtractor is used to extract audio data, image data, etc. from media data and convert it into a format that MediaCodec can process. MediaMuxer is used to encapsulate the encoded audio and video data into a multimedia container format, and it cooperates with MediaCodec to realize the output of encoded data.

[0046] Based on the media service framework provided by Android, the application layer application (application, APP) can call the above service components to implement data encoding and decoding functions.

[0047] The media service framework provided by Android has the following shortcomings:

[0048] 1. The media processing capabilities are split too finely, which is not conducive to vertical optimization. For example, when an application uses the encoding and decoding function, it needs to call different service components and communicate information between different service components, which is cumbersome.

[0049] 2. The application-oriented interface functions provided by MediaCodec are complex, unclearly named, and unfriendly to application developers. For example, MediaCodec contains a large number of interface functions, each of which is used to implement different functions, such as an interface function for storing image data in portable network graphic format (PNG), an interface function for storing image data in JPEG, and an interface function for storing audio and video data in audio video interleaved (avi). This requires application developers to specify the format in which the data is encoded, and to be familiar with the various interface functions provided by MediaCodec, so as to call the interface functions in a targeted manner to achieve the purpose.

[0050] 3. The media service framework fails to open the underlying differential capabilities, and the applications at the application layer cannot call the underlying differential capabilities through this media service framework. Specifically, the models or types of the processing chips used in different electronic devices may be different, that is, there are differences in hardware. The capabilities shared by most chips are called common capabilities, which can be called by the applications at the application layer through the media service framework. For the differential capabilities unique to each chip rather than shared, if the application developers are not familiar with these differential capabilities, the applications they develop cannot call these differential capabilities. When the differential capabilities include encoding and decoding capabilities, each application cannot call these encoding and decoding capabilities.

[0051] This application provides a software architecture of an electronic device, which can avoid the above deficiencies.

[0052] The electronic device provided by the embodiment of this application can be configured with an operating system (OS), and the OS can be one of etc.

[0053] Figure 2 shows the software structure of the OS of the electronic device provided by the embodiment of this application, and this software structure includes a unified media framework.

[0054] As Figure 2 shown, the electronic device can include the following four layers from top to bottom: the application layer, the framework layer, the service layer, and the kernel layer. The higher the layer, the more interactions with the user; the lower the layer, the more system capabilities it represents.

[0055] The application layer includes a series of user-oriented application programs (APPs), such as a recorder, a camera, a gallery, a video application, a music player software, a radio application, etc. Each of the above application programs can provide a user interface for the user to view and input user operations, so that the electronic device can respond to these user operations to generate new data or view existing data. For example, the audio player software can be used to play audio data, the video application can be used to play video data, the recorder can be used to collect audio data and generate audio data, and the camera application can be used to collect audio-visual data and generate audio-visual data. The electronic device will use encoding capabilities during the process of storing data, and will use decoding capabilities before editing or playing data.

[0056] The applications included in the application layer can be system applications or third-party applications. System applications include the applications configured to make the electronic device work properly and maintain the OS, and third-party applications can include the applications developed by other developers other than the electronic device manufacturer.

[0057] The framework layer, also known as the program framework layer, is used to provide application programming interfaces (APIs) and programming frameworks for the application programs in the application layer. The program framework layer may include some predefined functions. The framework layer may also include modules such as a window manager, a content provider, a view system, a phone manager, a resource manager, etc., which will not be elaborated here. The framework layer is used to connect the application layer and the service layer.

[0058] As Figure 2 shown, the framework layer may include function interfaces such as a camera function interface (PhotoMode), an editing function interface (EditMode), a storage function interface (StorageMode), a sharing function interface (ShareMode), a playback function interface (PlayMode), a file management function interface (FileManageMode), etc. These function interfaces can be directly called by the application programs in the application layer.

[0059] Exemplarily, the camera function interface can be an interface (PhotoModeKit) or a selector (PhotoModePicker). Among them, multiple shooting function interfaces can be encapsulated in PhotoModePicker. One shooting function interface can be used to start one shooting function, and the shooting function can refer to: shooting function, video recording function, night scene function, portrait function, large aperture function, time-lapse photography function, artificial intelligence (AI) scene recognition function, barcode scanning function, face recognition function, face interaction function, and so on. PhotoModePicker can select the corresponding interface from multiple shooting function interfaces according to the parameters passed in by the application.

[0060] For example, when the electronic device 100 starts portrait shooting through an application with a shooting function, that is, a camera application, the camera application can call PhotoModePicker and pass the parameters passed in by the application. For example, the parameter indicates that the camera used for shooting is a front camera. Then, through this parameter, PhotoModePicker can select the shooting function interface corresponding to the portrait function to implement the shooting function of the camera application.

[0061] It should be understood that other function interfaces such as EditMode and StorageMode are similar to the description of PhotoMode and will not be elaborated here.

[0062] It can be seen that the application framework layer can provide a unified interface for multiple applications in the application layer to call, enabling these multiple applications to achieve the same function. For example, for the system camera application and third-party applications, they can all call PhotoMode to implement the photo-taking function. Compared with different applications calling different interfaces, it solves the problem of chaotic interface calls during the operation of applications.

[0063] As Figure 2 shown, the application framework layer may further include: a media data interface (MediaDataKit). The media data interface can be used to be called by multiple function interfaces such as PhotoMode, EditMode, StorageMode, ShareMode, PlayMode, FileManageMode, etc., to control the acquisition, encoding, storage, sharing, decoding, and playing of media data.

[0064] Among them, whether to perform encoding or decoding on the media data can be determined by the function started by the application. For example, if the function started by the camera application is the camera function, the operations performed on the media data may include acquisition and encoding after acquisition. Another example is that if the function started by the gallery application is the editing function, the operations performed on the media data may include decoding. Another example is that if the function started by the video application is the playing function, the operations performed on the media data may include decoding and playing after decoding. It can be seen that the encoding or decoding of the media data can be considered to be executed following the function started by the user facing the application (such as the camera function, editing function, playing function, etc.), and the encoding function or decoding function does not directly present an entry to the user.

[0065] As Figure 2 shown, the application framework layer may further include: audio capability (Audio), video capability (Video), audio-video capability (Camera). The audio capability is an API that provides audio capabilities, and through calling the audio capability, audio-related encoding and decoding capabilities can be achieved. The video capability is an API that provides video capabilities, and through calling the video capability, video-related encoding and decoding capabilities can be achieved. The camera capability is an API that provides image processing capabilities, and through calling the camera capability, the audio-video encoding and decoding capabilities provided by the OS can be obtained. In other implementation manners, the framework layer may further include other types of capabilities, such as text capabilities, etc., which are not limited here. The above several capabilities can be regarded as interfaces that can be called. These multiple capabilities can be called by the media data interface to implement the encoding of media data.

[0066] As Figure 2As shown in the figure, the framework layer may further include the client capabilities of an audio-video coder-decoder (AVcodec). The AVcodec client capabilities may be a single client capability or may include three client capabilities, such as an audio client capability for being called by the audio capability, a video client capability for being called by the video capability, and an audio-video client capability for being called by the audio-video capability, etc. The AVcodec client capabilities are used to call the AVcodec server capabilities in the service layer downwards.

[0067] As Figure 2 As shown in the figure, the service layer provides one or more service layer interfaces, and the service layer interfaces may encapsulate one or more kernel layer interfaces provided by the kernel layer. Specifically, the service layer may include the following modules: multimedia service, audio service, video service, camera service, AVcodec server capabilities in the service layer, other services, etc. The multimedia service interfaces upwards to each capability, such as the audio capability, the video capability, etc., and is used to provide the business logics of the above capabilities. The audio service interfaces upwards to the audio capability and is used to provide the business logic of the audio capability. The video service interfaces upwards to the video capability and is used to provide the business logic of the video capability. The camera service interfaces upwards to the camera capability and is used to provide the business logic of the camera capability. It can be seen that a capability can be implemented with the support of multiple service modules in the service layer. The above several service modules can be regarded as service layer interfaces.

[0068] The AVcodec client capabilities in the framework layer and the AVcodec server capabilities in the service layer are used to jointly provide the encoding and decoding processing capabilities for media data such as audio, video, and audio-video. Communication between the AVcodec client capabilities and the AVcodec server capabilities may be based on an inter-process communication (IPC) mechanism. Communication between the AVcodec client capabilities and the AVcodec server capabilities adopts a client-server mechanism. The AVcodec server capabilities may encapsulate different encoding and decoding sets and are used to provide corresponding encoding and decoding processing capabilities for different data types. The encoding and decoding sets include an encoding set and a decoding set. The encoding set includes encoding methods and algorithms, and the decoding set includes decoding methods and algorithms. In other words, the AVcodec server capabilities include multiple atomized capabilities for audio encoding, multiple atomized capabilities for video encoding, and multiple atomized capabilities for audio-video encoding. The AVcodec server capabilities can be regarded as an encoding and decoding register for storing encoded data and decoded data.

[0069] The capabilities of the AVcodec server can be further divided into: AVcodec management, AVcodec sub-services, and AVcodec post-processing.

[0070] AVcodec management provides hardware and software basic interfaces. The hardware basic interface can be used to call the driver of the hardware coder-decoder (Hcodec) to drive the hardware coder-decoder to work. The software basic interface can be used to call the driver of the software coder-decoder to drive the hardware coder-decoder to work.

[0071] The AVcodec sub-services can be called by AVcodec management to implement the basic logic of encoding and decoding.

[0072] AVcodec post-processing is used to implement the post-processing capabilities of encoding and decoding not provided by the AVcodec sub-services.

[0073] The capabilities of the AVCodec server include a structure for storing codec information. The capabilities of the AVCodec server can include the following variables:

[0074] const char* name: The name of the codec, which is short;

[0075] const char* long_name: The full name of the codec, which is relatively long;

[0076] enum AVMediaType type: Specifies the type of media data, whether it is video, audio, or subtitle;

[0077] enum AVCodecID id: The encoding / decoding ID, which is unique;

[0078] const AVRational* supported_framerates: The supported frame rates (only for video);

[0079] const enum AVPixelFormat* pix_fmts: The supported pixel formats (only for video);

[0080] const int* supported_samplerates: The supported sample rates (only for audio);

[0081] const enum AVSampleFormat* sample_fmts: The supported sample formats (only for audio);

[0082] const uint64_t* channel_layouts: The supported number of channels (for audio only);

[0083] int priv_data_size: The size of the private data.

[0084] Generally, the processing flow of the AVcodec server capabilities includes: 1. Register all codecs: av_register_all(); 2. Declare a pointer of type AVCodec, such as AVCodec* first_c; 3. Call the av_codec_next() function to obtain a pointer to the next codec in the linked list. By repeating this process, information about all codecs can be obtained. Note that to obtain a pointer to the first codec, the parameter of this function needs to be set to NULL.

[0085] The kernel layer is the layer between hardware and software. As Figure 2 shown, the kernel layer may include the following modules: microphone driver, camera driver, driver for the hardware coder-decoder (Hcodec), driver for the software coder-decoder, algorithms, etc. Among them, the hardware coder-decoder can be implemented as a chip or other form of hardware. The algorithms may include audio codec algorithms, may also include video codec algorithms, as well as codec algorithms for other data types, etc.

[0086] As the connection layer between software and hardware, the kernel layer can be compatible with different hardware devices. The models or types of processing chips used in different electronic devices may be different, that is, there are differences in hardware. In the electronic device OS provided in the embodiments of the present application, the kernel layer can provide one or more kernel layer interfaces, and these one or more kernel layer interfaces encapsulate underlying capabilities, and these underlying capabilities include codec capabilities. These underlying capabilities specifically include two types of capabilities: basic capabilities and differentiated capabilities. Basic capabilities refer to the capabilities common to most platforms, and differentiated capabilities refer to the differentiated capabilities unique to each platform rather than shared. Here, the platform refers to hardware devices such as chips. The kernel layer can provide driver programs and hardware access interfaces for calling basic capabilities, and can also provide driver programs and hardware access interfaces for calling differentiated capabilities, etc. In this way, no matter which chips the electronic device is loaded with, the OS of the electronic device is ready to call the capabilities of these chips, achieving compatibility. The underlying capabilities encapsulated by the above-mentioned kernel layer interfaces include codec capabilities. The kernel layer implements a compatible southbound ecological interface design, provides southbound interfaces for docking hardware devices, and simplifies the docking cost.

[0087] Figure 2 The functions of the provided modules have been clearly introduced above, and their names do not constitute a limitation. The above-mentioned modules can also be called other names.

[0088] Figure 2 The software architecture of the electronic device shown may also include more or fewer modules, which are not limited here. For example, the application layer may further include software for generating and displaying text data such as a notepad and a reader, the framework layer may further include a text kit, and the service layer may further include a text service and a codec service for text, etc.

[0089] This application also provides a unified media framework, which may include several above-mentioned modules in the framework layer and is used to provide unified media data processing capabilities for applications in the application layer. The unified media framework of this application can be implemented as a system capability, or as a software development kit (SDK), a resident service, a binary shared object (SO) file, etc. Among them, the resident service can be regarded as a resident process and is always on after the electronic device is powered on.

[0090] Figure 2 The software architecture of the electronic device shown is a new codec architecture, which realizes the four-layer architecture implementation of the codec function.

[0091] Based on Figure 2 the codec architecture shown, the calling process of each module in this codec architecture is introduced below.

[0092] When the electronic device 100 detects that the user operates to take a photo through the camera application, the gallery application can call the PhotoMode interface, and the PhotoMode interface then calls the MediaDataKit. The MediaDataKit determines that the operation to be performed on the media data is "encoding" and "storing" according to PhotoMode, and can also determine the type of the media data (such as audio type, video type, audio-video type, etc.) according to the information of the media data passed in by the camera application. After that, the MediaDataKit selects and calls the encoding capability corresponding to this media data type, and this encoding capability then calls the corresponding AVcodec client capability, and the AVcodec client capability then calls the AVcodec server capability. The AVcodec server capability calls the corresponding atomicized capabilities according to the passed-in parameters, such as audio encoding atomicized capability, video encoding atomicized capability, audio-video encoding atomicized capability, etc. After the encoding is completed, the MediaDataKit then calls the storage capability to realize the storage of the photo.

[0093] When the electronic device 100 detects that the user operates on editing a photo through the gallery application, the gallery application can call the EditMode interface, and the EditMode interface then calls the MediaDataKit. The MediaDataKit determines that the operations to be performed on the media data are "decoding" and "editing" according to EditMode, and can also determine the type of the media data (such as audio type, video type, audio-video type, etc.) based on the information of the media data passed in by the gallery application. After that, the MediaDataKit selects and calls the decoding capability corresponding to this media data type, and this decoding capability then calls the corresponding AVcodec client capability, and the AVcodec client capability then calls the AVcodec server capability. The AVcodec server capability calls the corresponding atomized capabilities according to the passed-in parameters, such as audio decoding atomized capability, video decoding atomized capability, audio-video decoding atomized capability, etc. After completing the decoding, the MediaDataKit then calls the editing capability to implement the editing of the photo.

[0094] When the electronic device 100 detects that the user operates on viewing a video through the gallery application, the gallery application can call the PlayMode interface, and the PlayMode interface then calls the MediaDataKit. The MediaDataKit determines that the operations to be performed on the media data include "decoding" and "playing" according to PlayMode, and can also determine the type of the media data (such as audio type, video type, audio-video type, etc.) based on the information of the media data passed in by the gallery application. After that, the MediaDataKit selects and calls the decoding capability corresponding to this media data type, and this decoding capability then calls the corresponding AVcodec client capability, and the AVcodec client capability then calls the AVcodec server capability. The AVcodec server capability calls the corresponding atomized capabilities according to the passed-in parameters, such as audio decoding atomized capability, video decoding atomized capability, audio-video decoding atomized capability, etc. After completing the decoding, the MediaDataKit then calls the playing capability to implement the playing of the video.

[0095] The parameters passed down by the application in the application layer can be transmitted layer by layer to the AVcodec server capability, facilitating the AVcodec server capability to call the corresponding atomized capabilities according to these parameters.

[0096] In some embodiments, after the XX Mode interface is called, it can bypass the MediaDataKit and directly call other framework layer capabilities below the MediaDataKit to implement the function of application startup.

[0097] In some embodiments, the XXMode interface can be a kit or a picker. When the XXMode interface is implemented as a picker, the MediaDataKit can also be correspondingly subdivided into multiple interfaces with more detailed functions, so that the XXMode interface can select a suitable interface from the multiple interfaces for invocation.

[0098] After the AVcodec server - side capabilities are invoked, it will continue to call the interfaces of the kernel layer at a lower level to implement the corresponding encoding and decoding functions.

[0099] Among them, Figure 2 Any function interface provided by the middle framework layer can be regarded as XXMode. For example, the camera function interface (PhotoMode), the editing function interface (EditMode), the storage function interface (StorageMode), the sharing function interface (ShareMode), the playback function interface (PlayMode), the file management function interface (FileManageMode), etc. The MediaDataKit is the media data interface.

[0100] The application can detect the user operation and trigger the start of a certain function, which can be used to perform one or more of the processes of collecting, encoding, storing, sharing, decoding, and playing a certain media data.

[0101] Then, after the application detects the user operation, it can call the XXMode interface. For example, if the camera application detects a photo - taking operation, the XXMode interface can be the PhotoMode interface; if the video application detects a video - playing operation, the XXMode interface can be the PlayMode interface.

[0102] Among them, the form of the XXMode interface is XXMode(DeviceName, AppName, key, value, extend). The XXMode interface includes the following parameters: the name of the device that calls the XXMode interface (deviceName), the name of the application that calls the XXMode interface (appName), the indication information of the parameters required for encoding and decoding (key), the content of the parameters required for encoding and decoding (value), and the extended parameter (extend).

[0103] The device that calls the XXMode interface can be the local device or a remote device connected to the local device. The application that calls the XXMode interface can be an application on the local device or an application on the remote device. The parameters required for encoding and decoding may include the parameters required for hierarchical encoding, frame rate, bit rate, resolution, transcoding method, etc. The indication information of the parameters may include parameter name, identifier, number, index, etc., and the parameter content may refer to the specific parameter value. Among them, device name, application name, and extended parameters are optional. When different XXMode interfaces are called, the application that calls the XXMode interface can pass different parameters to the XXMode interface. A key and a value form a key-value pair, and an XXMode interface may include one or more key-value pairs. In some embodiments, parameters such as deviceName, appName, key, and value can be determined by an application in the application layer and then transmitted to the corresponding XXMode interface in the framework layer; in other embodiments, if the application in the application layer does not transmit the above parameters, the OS of the electronic device can also fill them with default parameters. For example, the deviceName is default filled with the name of the local device, the appName is filled with the name of the application that calls XXMode, and the default key and value are filled into the XXMode interface to implement the encoding and decoding functions corresponding to the XXMode interface.

[0104] The Media Data Interface (MediaDataKit) can be implemented as MediaDataKit(type, key1, value1, key2, value2, extend). It includes the following parameters:

[0105] Type (type), indicating the type to which the operation performed by the current electronic device belongs. The type may include, but is not limited to, the following: general type (normal), acquisition, encoding, storage, sharing, decoding, playing, other, etc. The type can be determined according to the previous interface that calls the MediaDataKit interface. For example, if the previous interface is the PhotoMode interface, the electronic device can determine that encoding needs to be performed currently, so it is regarded as the corresponding encoding type; for example, if the previous interface is the EditMode interface, the electronic device can determine that decoding needs to be performed currently, so it is regarded as the corresponding decoding type; for example, if the previous interface is the PlayMode interface, the electronic device can determine that decoding needs to be performed currently, so it is regarded as the corresponding decoding type.

[0106] The indication parameter (key) of the function point is used to indicate the function point to be started. Each type can include multiple different function points. For example, for encoding type or decoding type, there may be different encoding methods or decoding methods. The indication parameter (key) of the function point is the key in the previous interface that calls the MediaDataKit interface. For example, if the previous interface is the XXMode interface, the indication parameter (key) of the function point is the key in the XXMode interface.

[0107] The value required by the function point is used to indicate the parameter value required to implement the function point. The value required by the function point is the value of the previous interface that calls the MediaDataKit interface.

[0108] The indication parameter (key) of a function point and the value (value) required by the function point are in a corresponding relationship, which can be one pair or multiple pairs. The value (value) corresponding to the indication parameter (key) of a function point can be one or multiple.

[0109] The extended field (extend) can be used to build pre-embedded capabilities or differentiated capabilities.

[0110] based on Figure 2 The software architecture is shown, and the encoding and decoding method provided by the embodiment of the present application is introduced below.

[0111] The coding and decoding methods provided in the embodiments of the present application are divided into two types, one is hardware coding and decoding, and the other is software coding and decoding. Hardware coding and decoding is implemented based on hardware codecs, and software coding and decoding is implemented based on software codecs. The following is a detailed introduction.

[0112] Hardware Codec

[0113] Figure 3 The hardware encoding and decoding process provided by the embodiment of the present application is exemplified.

[0114] like Figure 3 As shown, the application layer of the application layer calls down to the AVcodec client capability of the framework layer. During this period, the application layer passes down the parameters required for encoding and decoding.

[0115] AVcodec client capabilities can pass received parameters to AVcodec server capabilities. AVcodec client capabilities can encapsulate multiple atomic capabilities for encoding and multiple atomic capabilities for decoding. An atomic capability is used to implement an encoding function or a decoding function. The encoding function here includes encoding-related capabilities such as encoding algorithms and encoding formats, and the decoding function includes decoding-related capabilities such as decoding algorithms and decoding formats.

[0116] Afterwards, the AVcodec server capability runs the corresponding encoding atomization capability or decoding atomization capability, and encapsulates the corresponding kernel layer interface through these encoding atomization capabilities or decoding atomization capabilities. These kernel layer interfaces may include a hardware device interface (HDI) (Hcodec HDI) for calling the hardware codec. Afterwards, the electronic device can call the hardware codec (Hcodec) and the corresponding codec algorithm, such as audio codec algorithm, video codec algorithm, etc. through the HDI.

[0117] pass Figure 3 It can be seen from the hardware encoding and decoding process shown that in the embodiment of the present application, the application of the application layer only needs to call the AVcodec client capability layer by layer, and the AVcodec server capability can call the corresponding atomic capability according to the incoming information to implement the encoding and decoding function required by the application. The OS actively provides unified encoding and decoding atomic capabilities in the AVcodec server capability, and there is no need for the application to directly call these atomic capabilities. The OS can call the appropriate XXMode interface for the application to implement the encoding and decoding function required by the application. This calling method is more friendly to the application, shielding the complex encoding and decoding operations and algorithms, and can simplify the programming operations of the application developer. Therefore, the AVcodec client capability and AVcodec server capability provided by this application are convenient for each application at the application layer to call, and can solve the problem of being unfriendly to the northbound access APP.

[0118] Hardware codecs can be responsible for the following management tasks: state machine management, thread management, and memory management. State machine management refers to the management and control of the state transition of objects or systems, thread management refers to the dynamic creation, extinction, suspension, and resumption of threads, and memory management refers to the allocation, recovery, and protection of memory.

[0119] In the embodiment of the present application, the hardware codec can have the following two improvements when working: 1. Reduce the copy of compressed frames.

[0120] Reducing compressed frame copies is applied to scenarios where different applications process the same data, for example, application A stores data and application B plays data.

[0121] Figure 4A An encoding and decoding process based on the media service framework provided by Android is exemplified.

[0122] like Figure 4AAs shown, data copying generally involves three layers: the application layer, the framework layer, and the system extension layer. Both the application layer and the framework layer belong to the operating system (OS) of the electronic device, and the system extension layer is provided by the hardware of the electronic device. The application layer may include applications. In the media service framework provided by Android, applications usually cannot directly schedule the underlying hardware (such as a hardware codec), but first call the framework layer of the OS, and then the framework layer of the OS calls the underlying hardware required by the third-party application. Therefore, as Figure 4A shown, there are the following two data copies:

[0123] The first copy is to copy data from an application in the application layer (such as a system player or a third-party application) to the buffer of the memory block provided by MediaCodec in the system framework layer. The buffer of the memory block provided by MediaCodec is generally virtual memory allocated through dynamic memory allocation (memory allocation, malloc), which neither supports cross-process sharing nor access by the hardware codec.

[0124] The second copy is to copy data from the memory block buffer in the system framework layer to the memory block buffer in the system extension layer provided by the hardware. The memory block buffer provided by the system extension layer is usually allocated by the function dma_buf, which supports cross-process sharing and access by the hardware codec. Therefore, after the second copy, the hardware codec can access the data in the memory block buffer in the system extension layer, thereby performing encoding processing on the data to obtain the corresponding data.

[0125] When Application B needs to play the data obtained by the above encoding, it first asks the framework layer for the data, and then the framework layer obtains the corresponding data from the system extension layer and decodes the data through the hardware codec. The decoded data can be returned to Application B for playback.

[0126] Since neither Application A nor Application B can directly schedule the underlying hardware and can only interact with the framework layer, based on the media service framework provided by Android, two copies need to be performed during encoding.

[0127] Figure 4B An example shows the encoding and decoding process based on the software framework provided by this application.

[0128] The data copying method provided by this application only involves two layers: the application layer and the system extension layer. In the solution provided in the embodiments of this application, the OS and the underlying hardware cooperate to break the barrier between the framework layer and the system extension layer. Therefore, it supports directly copying the data in the application layer to the memory block cache in the system extension layer.

[0129] When storing data, Application A can directly copy the data to the framework layer through the OS. Since both the framework layer and the system extension layer are controlled by the OS, Application A doesn't need to care about how to copy the data to the system extension layer provided by the hardware. Just copying the data to the framework layer can be regarded as having copied the data to the system extension layer. After the OS of this application obtains the data copied by Application A in the framework layer, it can copy the data to the system extension layer by itself without Application A having to do so. Then, the hardware codec can perform an encoding operation on the data to obtain the corresponding data. After that, when Application B wants to play the data, it only needs to pass the identifier of the data to the OS. The OS can then obtain the data from the system extension layer and perform decoding on the data through the hardware codec to obtain the decoded data, and the decoded data can be returned to Application B for playback.

[0130] Comparison Figure 4A and Figure 4B In Figure 4A Application A and Application B need to pay attention to the mapping relationship between the buffer in the framework layer and the buffer in the system extension layer in order to copy the data of Application A to the system extension layer for the hardware codec to encode or decode. In Figure 4B Application A and Application B only need to interact with the framework layer of the OS. The OS itself copies the data in the framework layer to the system extension layer for the hardware codec to encode or decode, and Application A and Application B don't need to pay attention to the mapping relationship between the buffer in the framework layer and the buffer in the system extension layer. In this way, it is equivalent that the OS shields the complex mapping relationship between the framework layer and the system extension layer, enabling the applications in the application layer not to care about the matters of the system extension layer and reducing the operations performed by the applications in the application layer.

[0131] Comparison Figure 4B and Figure 4A It can be seen from the data copying methods of Figure 4B The method shown in

[0132] 2. Optimize the display of the view.

[0133] Normally, the data output after the hardware codec performs a decoding operation on the data is stored in the buffer of the resource pool. The resource pool may include multiple buffers, and each buffer is bound or corresponds to a view. An electronic device usually includes a display screen, and the display screen can be used to display the picture provided after the synthesis of multiple views. The buffer can be regarded as an area for storing data, and the data can be output after being stored in the buffer for a period of time.

[0134] Figure 5A Exemplarily shows the corresponding relationship between the buffers and views in the resource pool. As Figure 5AAs shown, the buffer and view are strongly coupled one-to-one, and buffer 1 (buffer1) corresponds to view Figure 2 (view1), buffer 2 (buffer2) corresponds to view Figure 3 (view2), buffer 3 (buffer3) corresponds to view 4 (view3). The data in the buffer is only output to the corresponding view. If a buffer is damaged or abnormal, the subsequent view corresponding to the buffer will not be able to obtain valid data streams.

[0135] The embodiment of the present application decouples the strong correspondence between the buffer and the view, and eliminates the dependence of the view on the corresponding buffer, so that when the view needs it, the electronic device can retrieve the data stream from the buffer of the resource pool as needed and output it to the corresponding view.

[0136] In some cases, the buffer may become corrupted or abnormal. Figure 5B This example shows the optimal display mode when the buffer is damaged or abnormal. Damage or abnormality may occur when the buffer memory is insufficient or the hardware fails. Figure 5B As shown, view1 corresponds to buffer1. If buffer1 is damaged or abnormal, the electronic device will make adjustments based on the actual situation.

[0137] like Figure 5B As shown, if buffer1 is damaged or abnormal in the short term, the electronic device can replace buffer1 with the buffer in the resource pool that is most similar to buffer1, and output the data stream in the buffer that is most similar to buffer1 to view1. After the data stream of buffer1 is restored, the electronic device can continue to output the data stream in buffer1 to view1. The similarity of buffers refers to the similarity of the contents of two buffers. The electronic device can determine whether the two are similar by traversing the binary coded contents of the two buffers for comparison.

[0138] like Figure 5B As shown, if buffer1 is damaged or abnormal for a long time, the electronic device can find a buffer that best matches view1 in other buffers other than buffer1 to replace buffer1, and output the data stream in the buffer that best matches view1 to view1. The matching degree between buffer and view can be determined by the application or user. The original buffer1 matches view1. After buffer1 is damaged or abnormal for a long time, the electronic device can find other buffers that best match view1 to replace buffer1.

[0139] Figure 5B The two processing methods shown can be combined for implementation. For example, when the electronic device initially detects that buffer1 is damaged or abnormal, it first outputs the data stream of the buffer that is most similar to buffer1 to view1. If buffer1 is not detected to recover after a period of time, it outputs the data stream of the buffer that best matches view1 to view1.

[0140] Through Figure 5B the method shown, the electronic device can try to find a suitable buffer to output to the view, avoiding scenarios such as a black screen or other user-unfriendly situations caused by the view having no input data.

[0141] Software encoding and decoding

[0142] Figure 6 Exemplarily shows the software encoding and decoding process provided by the embodiments of the present application.

[0143] Figure 6 The software encoding and decoding process shown and Figure 3 the hardware encoding and decoding process shown are different in that in Figure 6 it, the AVcodec server capabilities run the encoding atomization capabilities or decoding atomization capabilities corresponding to the parameters passed in by the application, call the corresponding kernel layer interfaces through these encoding atomization capabilities or decoding atomization capabilities, and then call the software codec and the corresponding encoding and decoding algorithms, such as audio encoding and decoding algorithms, video encoding and decoding algorithms, etc., through running the kernel layer interfaces. The encoding atomization capabilities or decoding atomization capabilities here are different from the encoding atomization capabilities or decoding atomization capabilities used in the hardware encoding and decoding process. It can be seen that Figure 6 the encoding and decoding functions are implemented through a software codec, Figure 3 and the encoding and decoding functions are implemented through a hardware codec.

[0144] Such as Figure 6 shown, the software codec is also responsible for the following management tasks: state machine management, thread management, and memory management. The definitions of the above management tasks can refer to the relevant introduction to the hardware codec in the previous text Figure 3 for reference.

[0145] Whether it is hardware encoding and decoding or software encoding and decoding, through the encoding and decoding process provided by this application, the AVcodec client capabilities in the framework layer and the AVcodec server capabilities in the service layer interact in the communication mode of client capabilities - server capabilities. The applications in the application layer only need to call the XXMode interface provided by the OS in the framework layer, and then call the AVcodec client capabilities and AVcodec server capabilities layer by layer to obtain the encoding and decoding capabilities. Compared with the numerous interface functions in MediaCodec in the media service framework provided by Android, the XXMode interface provided in the framework layer has a simple and single function and a clear naming. Therefore, the XXMode interface in the framework layer provided by this application is convenient for various applications in the application layer to call and can solve the problem of being unfriendly to the northbound access APPs.

[0146] Transcoding scheme

[0147] Transcoding is the process of converting encoded data from one encoding format to another. The purpose of transcoding is to solve the problem of incompatible encoding formats. Different encoding methods will generate data in different formats. When an electronic device wants to open a certain data, but the electronic device does not support the data in the current format, the electronic device cannot open the data. In such a case, the electronic device needs to convert the data into a format that it supports in order to continue opening the data. Common transcoding operations may include, but are not limited to: converting MP3 audio to advanced audio coding (AAC) format, converting H.264 video to H.265 format, converting PNG images to JPEG format, etc.

[0148] Figure 7 Exemplarily shows a transcoding process provided by an embodiment of this application. As Figure 7 shown, this process may include the following steps:

[0149] 1. The electronic device determines whether each application in the current device supports the format of the data to be opened, or determines whether the currently launched application supports the format of the data to be opened.

[0150] 2. If it supports, directly call the capabilities supported by the application to open the data, that is, decode the data and play the decoded data.

[0151] 3. If it does not support, determine the data type. The data type may include, for example, text, picture, audio, video, audio-video, etc.

[0152] 4. According to the data type, determine the corresponding transcoding algorithm. Different data types can correspond to different transcoding algorithms.

[0153] 5. Load the transcoding algorithm corresponding to the data type from the transcoding algorithm library.

[0154] 6. Call the conversion interface provided by the AVcodec management, and use the transcoding algorithm loaded in the transcoding algorithm library to perform transcoding operations on the data.

[0155] 7. Save the transcoded data, and then call the capabilities supported by the application to open the transcoded data.

[0156] Through Figure 7 The transcoding process shown, the electronic device can convert the data from a format not supported by the application to a supported format as much as possible, so that the application can open the data.

[0157] In Figure 7 In the transcoding process shown, the transcoding rules involved in steps 4 - 6 are the key points of the transcoding scheme provided by this application. The transcoding rules provided by this application can include the following two types: same-type transcoding and different-type transcoding.

[0158] Same-type transcoding means converting the data into another format that is supported by the electronic device, belongs to the same type as the data, and is similar or close to the current format. For example, for picture data, the original jpg format can be converted to the png format. Another example is that for audio data, the original MP3 format can be converted to the AAC format. Adopting the same-type transcoding rule can ensure that the user experience does not decline.

[0159] Different-type transcoding means that when no other format that is supported by the electronic device, belongs to the same type as the data, and is similar or close to the current format can be found, look for other formats that belong to different types from the current data for transcoding. Different-type transcoding can be further divided into the following two methods: upward classification and content conversion.

[0160] Upward classification means looking for a simpler type at the next level of the type to which the current data belongs, and converting the data into the format of that simpler type. For example, for picture data, it can be converted into text-type formats such as txt, rtf, etc. Another example is that for video data, it can be converted into picture-type formats such as PNG, JPEG, etc.

[0161] Content conversion means trying to find a format that the electronic device can open and converting the current data into that format. The format converted by the content conversion method may not be the optimal format for opening the data, but it can ensure the normal decoding of the media data.

[0162] Figure 8 Exemplarily shows a transcoding process taking a video in high-dynamic range (HDR) vivid format as an example. As Figure 8As shown in the figure, the electronic device first starts the video transcoding service of the application. Through this video transcoding service, it determines whether the video to be opened currently is in the HDR vivid format. If so, it calls the video encoding and decoding capabilities or function interfaces of the framework layer. After that, the video framework interface layer can obtain the HDR vivid format video, and obtain an 8-bit (bit) bitstream and dynamic metadata based on this video. Then, it first decodes the video using the hardware decoding method. If it is successful, it ends. If it is not successful, it performs software decoding on the frame queue of each bitstream. Among them, software decoding depends on the transcoding capabilities provided by the transcoding algorithm library.

[0163] The transcoding algorithm library provides two transcoding capabilities: directional conversion and optimization conversion. Directional conversion means that the data types before transcoding (such as type A) and after transcoding (such as type B) are known in advance, and the transcoding algorithm is formulated according to the encoding and decoding algorithms of type A and type B. Optimization conversion means that when the data types before and after transcoding are not clear, the transcoding algorithm can be matched according to the bitstream. In the optimization conversion, there are two ways to match the transcoding algorithm according to the bitstream: 1. Match the transcoding algorithm for the overall bitstream, that is, match the binary bitstream one by one. 2. Divide the overall bitstream into blocks, match the transcoding algorithm for each block of bitstream, and then perform weighted calculation on each block of bitstream to obtain the final transcoding algorithm.

[0164] Through the transcoding method provided by the embodiments of the present application, the electronic device can try to transcode and play the data, so that the user can know the information contained in the data, and avoid the situation where the user cannot know the data information such as the electronic device going black screen or reporting an error.

[0165] Multi-channel picture quality improvement

[0166] The bit rate is a parameter in the encoding and decoding process, which defines the bit rate code rate used by the electronic device when playing content on the display screen. The higher the bit rate, the clearer the content displayed by the electronic device.

[0167] In the screen mirroring scenario, an electronic device can receive the screen mirroring content from other electronic devices. This electronic device can receive multiple screen mirroring contents from another electronic device, or can receive the screen mirroring contents from multiple other electronic devices. Different screen mirroring contents can be displayed in different areas of the display screen. Refer to Figure 9 , Figure 9 Exemplarily shows a screen mirroring scenario. In Figure 9 the shown scenario, the mobile phone projects three interfaces onto the large screen or personal computer (PC) for display respectively.

[0168] In the embodiments of the present application, after the electronic device recognizes a multi-screen projection scenario (i.e., a scenario where multiple screen projection contents are projected into the electronic device for display), it can recognize the categories of each of the multiple screen projection contents. The categories of the screen projection contents may include: low-bitrate content, static content, and dynamic content. Among them, the low-bitrate content refers to non-high-definition content with a relatively low bitrate, and the bitrates of both the static content and the dynamic content are higher than that of the low-bitrate content. The static content refers to the content whose interface remains unchanged for a period of time, such as text, pictures, etc. The dynamic content refers to the content whose interface changes frequently for a period of time, such as videos, etc.

[0169] In a multi-screen projection scenario, the electronic device can allocate the overall bitrate to multiple screen projection contents. The allocation principle is as follows: the bitrates of the three types of low-bitrate content, static content, and dynamic content decrease in sequence, and the sum of the bitrates of all screen projection contents is equal to the overall bitrate. The overall bitrate depends on the software and hardware capabilities of the electronic device. For example, if the overall bitrate of a large screen is 10, and the screen projection contents sent by other devices include three, one low-bitrate content, one static content, and one dynamic content, then the large screen allocates a bitrate of 2 to the low-bitrate content, a bitrate of 3 to the static content, and a bitrate of 5 to the dynamic content.

[0170] In some embodiments, the electronic device can also adjust the bitrates of each screen projection content according to the screen focus. The screen focus is the location of the user's current touch point and can be considered as the content that the user is concerned about. Therefore, the bitrate of the screen projection content where the screen focus is located can be increased to improve the display quality of the screen content, thereby giving the user a good usage experience. In some embodiments, the electronic device can adjust the bitrate of the screen projection content where the screen focus is located to the highest bitrate among all screen projection contents. For example, assuming that the electronic device initially allocates bitrates of 2, 3, and 5 to the low-bitrate content, static content, and dynamic content in sequence, when the electronic device detects that the screen focus is on the low-bitrate content, the electronic device can adjust the bitrates of the low-bitrate content, static content, and dynamic content to 5, 2, and 3 in sequence, so that the user can see high-quality low-bitrate content.

[0171] Hierarchical coding

[0172] When the electronic device encodes an audio-video stream, in order to avoid audio-video asynchronization, the audio stream and the video stream are usually subjected to unified encoding and decoding processing. However, due to different encoding and decoding algorithms for audio and video, this will result in excessive encoding and decoding operations and cannot play data while encoding and decoding.

[0173] In the embodiments of the present application, in order to improve the encoding efficiency of the streaming media, the audio stream and the video stream are processed separately, so that encoding optimization can be performed on the audio stream and the video stream respectively. In the embodiments of the present application, the video stream collected by the camera, the audio stream collected by the microphone, etc. can all be referred to as streaming media.

[0174] Reference Figure 10 , the processing procedure of the audio - video stream may include the following steps:

[0175] 1. Data acquisition. For example, the electronic device can start the camera application, activate the camera and microphone, and acquire video and audio.

[0176] 2. Frame processing. After the electronic device acquires the video, it performs video frame processing, that is, replaces the new video frame in the surface of the camera application. After the electronic device acquires the audio, it performs audio frame processing, such as filtering.

[0177] 3. Encoding. The electronic device encodes the processed video and audio respectively. For example, it uses a VideoEncoder to encode the video and an AudioEncoder to encode the audio.

[0178] 4. Encapsulation. After encoding, the electronic device can encapsulate the audio and video together, that is, encapsulate multiple frames of data into a MultimediaContainer. In some embodiments, the electronic device can also encapsulate the metadata for audio - video synchronization together.

[0179] 5. Decapsulation. When the electronic device receives an operation for viewing the video in the gallery application, the encapsulated multimedia container in the electronic device is transmitted from the camera application to the gallery application, and the gallery application performs decapsulation. Decapsulation is the reverse process of encapsulation.

[0180] 6. Obtaining frame data. After decapsulation, the electronic device can separate at least two paths of frame data, one path is video elementary stream (VideoES), and the other path is audio frame data (AudioES). In some embodiments, after encapsulating the data, the electronic device may generate a sub - stream for naming, and this sub - stream includes the names of each data segment or data names. Therefore, the obtained frame data may also include a path of subtitle data (Subtitle).

[0181] 7. Decoding. The electronic device can decode the video frame data and audio frame data respectively. For example, the electronic device can use a VideoDecoder to decode the video frame data and an AudioDecoder to decode the audio frame data. After decoding, data such as video and audio can be obtained.

[0182] From Figure 10It can be seen that the present application provides a separate parallel processing solution for video streams and audio streams. After obtaining the separate video stream and audio stream, the electronic device can perform encoding optimization operations on the streaming media. That is, Figure 10 the encoding in step 3 in Figure 10 can adopt the streaming media encoding optimization operation provided by the embodiments of the present application.

[0183] Figure 11A is an example of single-layer streaming media encoding. Single-layer means that the electronic device encodes the streaming media sequentially in time domain, and only encodes one data stream at a time. As Figure 11A shown, the overall data stream length of the streaming media is y, and y can be the duration for which the electronic device performs single-layer encoding on the data stream. As Figure 11A shown, after the encoding of the previous data stream is completed, the electronic device will start the encoding work of the next data stream. Therefore, the duration required for single-layer encoding is y.

[0184] Figure 11B is an example of hierarchical encoding of streaming media provided by the embodiments of the present application. Hierarchical encoding means that the electronic device encodes the streaming media sequentially in time domain, and can encode multiple data streams at a time. As Figure 11B shown, the overall data stream length of the streaming media is y, and y can be the duration for which the electronic device performs single-layer encoding on the data stream. As Figure 11B shown, the electronic device can divide the data stream into k (such as 4) parts, and the lengths of these k parts of the data stream are not limited and can be the same or different. After the encoding of the previous data stream starts, every timeset, the electronic device can start the encoding work of the next data stream. That is, the electronic device first encodes the first data stream, and after timeset from the start of encoding, encodes the second data stream, and after timeset from the start of encoding the second data stream, encodes the third data stream, and so on until the encoding work of all data streams is completed. Figure 11B In Figure 11B , the length of the square represents the encoding duration of different parts of the data stream, and the total duration of all squares added together is the overall data stream length y. Figure 11B In Figure 11B , different squares represent different time domains, and each square can respectively represent the 0th layer, 1st layer, 2nd layer, etc. of the time domain. Adopting Figure 11B the scheme shown in Figure 11B , the electronic device can encode multiple data streams at the same time. Adopting Figure 11B the hierarchical encoding scheme shown in Figure 11B , the duration required for the electronic device to complete the encoding of the data stream with length y is less than y, which is equivalent to saving the overall encoding time. According to experimental data, according to Figure 11B the encoding optimization method described in Figure 11B , the encoding efficiency of the streaming media can be increased by about 10%.

[0185] k and timeset can be pre-set. There are various ways to divide the data stream into k parts, which are not limited here. timeset is less than or equal to the length of each data stream among the k data streams, so as to ensure that two adjacent data streams can be encoded simultaneously within a certain period of time, saving the encoding duration.

[0186] Figure 11C This is an example of hierarchical encoding for streaming media provided by the embodiments of this application. Figure 11C Different from Figure 11B is that the electronic device performs time-domain superposition, that is, encodes two data streams simultaneously in the same time period. As Figure 11C shown, first divide the length y of the overall data stream into 2 segments evenly, and encode the data stream with a length of y / 2 in the first segment using Figure 11B 's method, and also encode the data stream with a length of y / 2 in the second segment using Figure 11B 's method. When encoding the two data streams with a length of y / 2, the same time-domain layer is used, that is, encode the two data streams with a length of y / 2 simultaneously at the same time. Using the Figure 11B shown scheme, the electronic device can encode more data streams at the same time. Using the Figure 11B shown hierarchical encoding scheme, the time required for the electronic device to complete the encoding of the data stream with a length of y is less than y, which is equivalent to saving the overall encoding time.

[0187] In some other embodiments, the electronic device can also divide the length of the overall data stream into more segments and encode more data streams in the same time period. This can further perform superposition in the time domain, saving the encoding duration, and also requires more encoding resources. Equivalently, the electronic device divides the data stream into n segments, and then divides each data stream into k parts, where n is greater than or equal to 1 and k is greater than or equal to 2. The length of the data stream is the time required for encoding the data stream; encode the n data streams in parallel. Among them, the encoding start times of the first data stream in the n data streams are the same. In each data stream, the encoding start time of the i-th data stream is earlier than that of the (i + 1)-th data stream by timeset, and timeset is less than or equal to the length of each data stream. n can be pre-set.

[0188] Based on Figure 10 shown precondition of separately encoding and decoding the audio stream and the video stream, the electronic device can separately obtain the audio stream and the video stream, and can use Figure 11B or Figure 11CThe hierarchical coding scheme shown is used to hierarchically encode an audio stream or a video stream to improve the coding efficiency and generate and save data more quickly. Similarly, during the decoding process, the electronic device can also adopt a similar hierarchical decoding scheme to hierarchically decode the audio stream or the video stream to improve the decoding efficiency and play the data more efficiently and with higher quality.

[0189] Each of the foregoing embodiments described in the embodiments of the present application can be combined and implemented.

[0190] Figure 12 It is a flowchart of the encoding and decoding method provided by the embodiments of the present application. As Figure 12 shown, the method may include the following steps:

[0191] S1201, the first electronic device runs a first application on a first operating system.

[0192] The first operating system is the OS mentioned in the foregoing embodiments of the present application, and reference can be made to Figure 2 and related descriptions.

[0193] The first operating system is used to run the first application. The first application is a third-party application, such as a camera application, a gallery application, etc.

[0194] The first operating system provides: a first interface, a media data interface, and a second capability. Among them, the first interface is one of the following provided by the framework layer: a camera function interface (PhotoMode), an editing function interface (EditMode), a storage function interface (StorageMode), a sharing function interface (ShareMode), a playback function interface (PlayMode), a file management function interface (FileManageMode). The media data interface is MediaDataKit.

[0195] The first application includes program code for calling the first interface. The first interface is used to start the first function provided by the first application. The first interface is used to call the media data interface, and the media data interface encapsulates the second capability. The second capability includes the ability to encode audio data, the ability to encode video data, and the ability to encode audio and video data provided by the framework layer.

[0196] Furthermore, the ability to decode audio data includes multiple atomic capabilities for audio decoding, the ability to encode video data includes multiple atomic capabilities for video encoding, and the ability to encode audio and video data includes multiple atomic capabilities for audio and video encoding.

[0197] In some embodiments:

[0198] The first client capability and the first server capability. The first client capability is used to transmit the indication information and parameter content of the parameters for performing audio encoding passed in by the first application to the first server capability. The first server capability is used to select one or more atomic capabilities from multiple audio encoding atomic capabilities according to the information passed in by the first application to encode the first media data;

[0199] The capability of performing encoding on video data includes: the second client capability and the first server capability. The second client capability is used to transmit the indication information and parameter content of the parameters for performing audio encoding passed in by the first application to the first server capability. The first server capability is used to select one or more atomic capabilities from multiple video encoding atomic capabilities according to the information passed in by the first application to encode the first media data;

[0200] The capability of performing encoding on audio-visual data includes: the third client capability and the first server capability. The third client capability is used to transmit the indication information and parameter content of the parameters for performing audio encoding passed in by the first application to the first server capability. The first server capability is used to select one or more atomic capabilities from multiple audio-visual encoding atomic capabilities according to the information passed in by the first application to encode the first media data;

[0201] Among them, the first client capability can be Figure 2 the audio client capability in Figure 2 the second client capability can be Figure 2 the video client capability in Figure 2 the third client capability can be

[0202] S1202, Detect a user operation to start the first function of the first application. The first function is used to trigger the encoding of the first media data.

[0203] This user operation can be, for example, an operation to start taking pictures, an operation to start recording videos, etc. The first media data can be data of types such as pictures, audio, video, audio-visual, text, etc.

[0204] S1203, Encode the first media data.

[0205] When the first electronic device encodes the first media data, it can call the interfaces and capabilities of the OS layer by layer from the application layer. For specific references, please refer to the relevant descriptions in the previous text Figure 2 and will not be elaborated here.

[0206] In some embodiments, the server capabilities may also return the encoding result to the corresponding client capabilities, and the client capabilities return it to the first application via the second capabilities, the media data interface, the first interface, etc. For example, the first server capabilities return the result of encoding the first media data to the first client capabilities or the second client capabilities or the third client capabilities, and the first client capabilities or the second client capabilities or the third client capabilities are also used to return the result of encoding the first media data to the first application.

[0207] S1203 may encode the first media data through the hardware codec or software codec of the first electronic device.

[0208] In some embodiments, the first electronic device may also use a hierarchical encoding scheme to encode the first media data. For details, reference can be made to Figures 11A - 11C the description.

[0209] S1204, run the second application on the first operating system.

[0210] The first operating system is also used to run the second application. The second application is a third-party application, such as a gallery application, a video application, etc.

[0211] The first operating system also provides: a second interface. Among them, the second interface is one of the following provided by the framework layer: camera function interface (PhotoMode), editing function interface (EditMode), storage function interface (StorageMode), sharing function interface (ShareMode), playback function interface (PlayMode), file management function interface (FileManageMode).

[0212] The first application includes program code for invoking the second interface. The second interface is used to start the fourth function provided by the second application. The second interface is used to call the media data interface, and the media data interface encapsulates the second capabilities. The second capabilities also include the capabilities provided by the framework layer for decoding audio data, decoding video data, and decoding audio-visual data.

[0213] Furthermore, the ability to decode audio data includes multiple atomic audio decoding capabilities, the ability to decode video data includes multiple atomic video decoding capabilities, and the ability to decode audio-visual data includes multiple atomic audio-visual decoding capabilities.

[0214] In some embodiments:

[0215] The first client capability and the first server capability. The first client capability is used to transmit the indication information and parameter content of the parameters for performing audio decoding passed in by the first application to the first server capability. The first server capability is used to select one or more atomized capabilities from multiple atomized audio decoding capabilities according to the information passed in by the first application to decode the first media data;

[0216] The capability to decode video data includes: the second client capability and the first server capability. The second client capability is used to transmit the indication information and parameter content of the parameters for performing audio decoding passed in by the first application to the first server capability. The first server capability is used to select one or more atomized capabilities from multiple atomized video decoding capabilities according to the information passed in by the first application to decode the first media data;

[0217] The capability to decode audio - video data includes: the third client capability and the first server capability. The third client capability is used to transmit the indication information and parameter content of the parameters for performing audio decoding passed in by the first application to the first server capability. The first server capability is used to select one or more atomized capabilities from multiple atomized audio - video decoding capabilities according to the information passed in by the first application to decode the first media data.

[0218] S1205, detect a user operation to start the fourth function of the first application, where the fourth function is used to trigger the decoding of the first media data.

[0219] This user operation can be, for example, an operation to start editing an image, start playing a video, start viewing a picture, etc. The first media data can be data of types such as pictures, audio, video, audio - video, text, etc.

[0220] S1206, decode the first media data.

[0221] When the first electronic device encodes the first media data, it can call the interfaces and capabilities of the OS layer by layer from the application layer. For specific references, please refer to the relevant descriptions in the previous text Figure 2 and will not be elaborated here.

[0222] In some embodiments, the server capability can also return the decoding result to the corresponding client capability, and the client capability returns it to the first application via the second capability, the media data interface, the first interface, etc. For example, the first server capability returns the result of decoding the first media data to the first client capability or the second client capability or the third client capability, and the first client capability or the second client capability or the third client capability is also used to return the result of decoding the first media data to the first application.

[0223] S1206 can decode the first media data through the hardware codec or software codec of the first electronic device.

[0224] When decoding the first media data through the hardware codec, the decoupling relationship between the buffer and the view can be decoupled. For example, the first electronic device can store the decoded data in the first buffer, and output the data in the first buffer to the first view displayed on the display screen of the first electronic device for data playback; if it is detected that the first buffer is damaged, the data in the second buffer is output to the first view for data playback.

[0225] Among them, if the duration of the damage of the first buffer is less than the first value, the second buffer is the buffer most similar to the first buffer among the buffers for storing decoded data; if the duration of the damage of the first buffer is greater than the first value, the second buffer is the buffer most matching the first view among the buffers for storing decoded data. The first value can be preset and will not be specifically limited here.

[0226] In some embodiments, if the first media data is in a format not supported by the atomization capability in the second capability, before S1206, the first electronic device converts the first media data from the first format to the second format, and then can decode the first media data in the second format.

[0227] The second format is a format supported by the atomization capability in the second capability.

[0228] Among them, there are two transcoding methods: 1. Same-type transcoding, for example, both the first format and the second format are formats of the first data type (such as text, picture, video, audio, or audio-video), and the first data type is the type of the first media data. 2. Upward classification transcoding. For example, the first format is a format of the first data type (such as video), and the second format is a format of the second data type (such as text), and the data volume of the data of the second data type is less than the data volume of the data of the first data type.

[0229] In some embodiments, the first electronic device can also adjust the bit rate of the screen mirroring content. Specifically, the first electronic device receives the screen mirroring content of the second electronic device and the screen mirroring content of the third electronic device, plays the screen mirroring content of the second electronic device at the first bit rate, and plays the screen mirroring content of the third electronic device at the second bit rate. The first bit rate is less than the second bit rate, and the sum of the first bit rate and the second bit rate is equal to the total bit rate of the first electronic device; among them, the screen mirroring content of the second electronic device is static content, and the screen mirroring content of the third electronic device is dynamic content; or, the screen mirroring content of the second electronic device is low-bit content, and the screen mirroring content of the third electronic device is dynamic content; or, the screen mirroring content of the second electronic device is low-bit content, and the screen mirroring content of the third electronic device is static content.

[0230] In some embodiments, if it is detected that the screen focus is on the screen mirroring content of the first electronic device, the screen mirroring content of the second electronic device is played at a third bitrate, and the screen mirroring content of the third electronic device is played at a fourth bitrate, where the third bitrate is greater than the fourth bitrate, and the sum of the third bitrate and the fourth bitrate is equal to the total bitrate of the first electronic device.

[0231] In some embodiments, the first electronic device can also perform hierarchical encoding on the first media data, and the logic of the hierarchical encoding can refer to Figure 11B and Figure 11C and the relevant written descriptions.

[0232] Refer to Figure 13 , Figure 13 , which is the hardware structure diagram of the electronic device 100 provided by the embodiment of the present application. The electronic device 100 can be the electronic device and the first electronic device mentioned above. The OS introduced above runs on the electronic device 100.

[0233] As Figure 13 shown, the electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a microphone 170C, a camera 193, a display screen 194, a hardware codec 195, etc.

[0234] It can be understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0235] The processor 110 may include one or more processing units. For example: the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0236] The controller can generate operation control signals according to the instruction operation code and the timing signal to complete the control of fetching instructions and executing instructions.

[0237] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can hold the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0238] The electronic device 100 implements the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, which is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change the display information.

[0239] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD). The display panel can also be made of an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniled, a microled, a micro-oled, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device may include 1 or N display screens 194, where N is a positive integer greater than 1.

[0240] The electronic device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.

[0241] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise and brightness of the image through algorithms. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be provided in the camera 193.

[0242] The camera 193 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard formats such as RGB and YUV. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0243] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0244] The video codec is used to compress or decompress digital videos. The electronic device 100 may support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0245] The NPU is a neural-network (NN) computing processor. By learning from the biological neural network structure, such as learning from the transmission function between human brain neurons, it can quickly process the input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.

[0246] The internal memory 121 may include one or more random access memories (RAM) and one or more non-volatile memories (NVM).

[0247] Random access memory may include static random-access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM, for example, the fifth generation of DDR SDRAM is generally called DDR5 SDRAM), etc.; non-volatile memory may include disk storage devices and flash memory.

[0248] Flash memory can be classified into NOR FLASH, NAND FLASH, 3D NAND FLASH, etc. according to the operating principle, and can be classified into single-level cell (SLC), multi-level cell (MLC), triple-level cell (TLC), quad-level cell (QLC), etc. according to the number of potential levels of storage units. According to the storage specification, it can include universal flash storage (UFS), embedded multi media Card (eMMC), etc.

[0249] Random access memory can be directly read and written by the processor 110, and can be used to store the operating system or executable programs of other running programs (such as machine instructions), and can also be used to store data of users and application programs, etc.

[0250] Non-volatile memory can also store executable programs and data of users and application programs, etc., and can be pre-loaded into random access memory for direct reading and writing by the processor 110.

[0251] The external memory interface 120 can be used to connect to an external non-volatile memory to expand the storage capacity of the electronic device 100. The external non-volatile memory communicates with the processor 110 through the external memory interface 120 to achieve the data storage function. For example, data such as music and video are saved in the external non-volatile memory.

[0252] The microphone 170C, also known as a "microphone" or "transmitter", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak close to the microphone 170C with their mouth to input the sound signal into the microphone 170C. The electronic device 100 may be provided with at least one microphone 170C. In some other embodiments, the electronic device 100 may be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 may also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.

[0253] The hardware codec 195 is a hardware device dedicated to encoding and decoding, and can be used for encoding and decoding of audio, video, pictures, text, audio-video and other types of data. The hardware codec 195 may be integrated in the CPU, GPU or a dedicated video processing unit (VPU). The number of hardware codecs 195 may be one or more. If the electronic device includes multiple hardware codecs 195, each codec 195 can be used for encoding and decoding different types of data.

[0254] In the embodiments of the present application, the internal memory 121 is used to store the computer program for implementing the encoding and decoding method provided in the embodiments of the present application, and the processor 110 is used to execute the computer program to implement the encoding and decoding method provided in the embodiments of the present application.

[0255] It should be understood that the steps in the above method embodiments can be completed by the integrated logic circuit of the hardware in the processor or the instructions in software form. The method steps disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor.

[0256] The present application also provides an electronic device, which may include a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the method executed by the electronic device in any of the above embodiments.

[0257] The present application also provides a chip system, including a processing circuit and an interface circuit. The interface circuit is used to receive computer instructions and transmit them to the processing circuit, and the processing circuit is used to run the computer instructions to implement the method executed by the electronic device in any of the above embodiments.

[0258] The present application also provides a chip system, which includes at least one processor for implementing the method executed by the electronic device in any of the above embodiments. In a possible design, the chip system further includes a memory for storing program instructions and data, and the memory is located inside or outside the processor.

[0259] The chip system may be composed of chips, or may include chips and other discrete devices.

[0260] Optionally, the processor in the chip system may be one or more. The processor may be implemented by hardware or by software. When implemented by hardware, the processor may be a logic circuit, an integrated circuit, etc. When implemented by software, the processor may be a general-purpose processor that implements by reading the software code stored in the memory.

[0261] Optionally, the memory in the chip system may also be one or more. The memory may be integrated with the processor or separately provided from the processor, and the embodiments of the present application do not limit this. Exemplarily, the memory may be a non-transitory processor, such as a read-only memory (ROM), which may be integrated with the processor on the same chip or separately provided on different chips. The embodiments of the present application do not specifically limit the type of the memory and the setting manner of the memory and the processor.

[0262] Exemplarily, the chip system may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processing unit (CPU), a network processor (NP), a digital signal processing circuit (DSP), a micro controller unit (MCU), a programmable logic device (PLD), or other integrated chips.

[0263] The present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method executed by the electronic device in any of the above embodiments is implemented.

[0264] The present application also provides a computer program product, including a computer program, which when executed by a processor, implements the method executed by the electronic device in any of the above embodiments.

[0265] The various embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0266] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state disk (SSD)), etc.

[0267] Those of ordinary skill in the art can understand all or part of the processes in the methods of the above embodiments. These processes can be completed by relevant hardware instructed by a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes: ROM or random access memory RAM, magnetic disk, or optical disk, etc., which can store program codes of various types.

[0268] In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B can represent A or B; "and / or" in the text is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0269] The terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0270] In summary, the above description is only for the embodiments of the technical solution of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made in accordance with the disclosure of the present application shall be included within the protection scope of the present application.

Claims

1. A coding method, characterized in that, The method is applied to a first electronic device on which a first operating system runs, and the method includes: Run a first application on the first operating system. The first application is installed on the first electronic device. The first operating system provides: a first interface, a media data interface, and a second capability. The first application is a third-party application and includes program code for calling the first interface. The first interface is used to call the media data interface, and the media data interface encapsulates the second capability. The second capability includes: the ability to perform encoding on audio data, the ability to perform encoding on video data, and the ability to perform encoding on audio-video data. Among them, the ability to perform encoding on audio data includes multiple atomic capabilities for audio encoding, the ability to perform encoding on video data includes multiple atomic capabilities for video encoding, and the ability to perform encoding on audio-video data includes multiple atomic capabilities for audio-video encoding; Detect a user operation to start a first function of the first application, where the first function is used to trigger encoding of first media data; Encode the first media data.

2. The method according to claim 1, wherein: The ability to perform encoding on audio data includes: a first client capability and a first server capability. The first client capability is used to transmit indication information and parameter content of parameters for performing audio encoding passed in by the first application to the first server capability. The first server capability is used to select one or more atomic capabilities for audio encoding from multiple atomic capabilities for audio encoding according to the information passed in by the first application to encode the first media data; The ability to perform encoding on video data includes: a second client capability and a first server capability. The second client capability is used to transmit indication information and parameter content of parameters for performing audio encoding passed in by the first application to the first server capability. The first server capability is used to select one or more atomic capabilities for video encoding from multiple atomic capabilities for video encoding according to the information passed in by the first application to encode the first media data; The ability to perform encoding on audio-video data includes: a third client capability and a first server capability. The third client capability is used to transmit indication information and parameter content of parameters for performing audio encoding passed in by the first application to the first server capability. The first server capability is used to select one or more atomic capabilities for audio-video encoding from multiple atomic capabilities for audio-video encoding according to the information passed in by the first application to encode the first media data.

3. The method according to claim 2, wherein: The first server capability is further used to return the encoding result of the first media data to the first client capability or the second client capability or the third client capability, and the first client capability or the second client capability or the third client capability is further used to return the encoding result of the first media data to the first application.

4. The method according to any one of claims 1-3, characterized in that, The method further includes: Run a second application on the first operating system. The second application is installed on the first operating system. The first operating system also provides: a second interface. The second application is a third-party application, and the second application includes program code for invoking the second interface. The second interface is used to start a fourth function provided by the second application. The second interface is used to invoke the media data interface, and the media data interface encapsulates the second capability. The second capability further includes: the ability to decode audio data, the ability to decode video data, and the ability to decode audio-video data. Among them, the ability to decode audio data includes multiple atomic audio decoding capabilities, the ability to decode video data includes multiple atomic video decoding capabilities, and the ability to decode audio-video data includes multiple atomic audio-video decoding capabilities; Detect a user operation to start the fourth function of the first application. The fourth function is used to trigger the decoding of the first media data; Decode the first media data.

5. The method according to claim 4, wherein, The ability to decode audio data includes: a first client capability and a first server capability. The first client capability is used to transmit the indication information and parameter content of the parameters for performing audio decoding passed in by the first application to the first server capability. The first server capability is used to select one or more atomic capabilities from multiple atomic audio decoding capabilities to decode the first media data according to the information passed in by the first application; The ability to decode video data includes: a second client capability and a first server capability. The second client capability is used to transmit the indication information and parameter content of the parameters for performing audio decoding passed in by the first application to the first server capability. The first server capability is used to select one or more atomic capabilities from multiple atomic video decoding capabilities to decode the first media data according to the information passed in by the first application; The ability to decode audio-video data includes: a third client capability and a first server capability. The third client capability is used to transmit the indication information and parameter content of the parameters for performing audio decoding passed in by the first application to the first server capability. The first server capability is used to select one or more atomic capabilities from multiple atomic audio-video decoding capabilities to decode the first media data according to the information passed in by the first application.

6. The method according to claim 5, wherein, The first server capability is further used to return the result of decoding the first media data to the first client capability or the second client capability or the third client capability. The first client capability or the second client capability or the third client capability is further used to return the result of decoding the first media data to the first application.

7. The method according to any one of claims 4-6, characterized in that, Decoding the first media data specifically includes: Decode the first media data through the hardware codec or software codec of the first electronic device.

8. The method according to claim 7, characterized in that, Decode the first media data through the hardware codec of the first electronic device, the method further comprising: Store the decoded data in a first buffer; Output the data in the first buffer to a first view displayed on the display screen of the first electronic device for completing data playback; Detect that the first buffer is damaged; Output the data in the second buffer to the first view for completing data playback.

9. The method according to claim 8, wherein if the duration of damage to the first buffer is less than a first value, the second buffer is the buffer most similar to the first buffer among the buffers for storing decoded data; if the duration of damage to the first buffer is greater than the first value, the second buffer is the buffer most matching the first view among the buffers for storing decoded data.

10. The method according to any one of claims 4-9, characterized in that, Decoding the first media data specifically includes: Convert the first media data from a first format to a second format, wherein the atomization capability in the second capability does not support the first format and supports the second format; Decode the first media data in the second format.

11. The method according to claim 10, wherein both the first format and the second format are formats of a first data type, and the first data type is the type of the first media data; or the first format is a format of the first data type, and the second format is a format of a second data type, and the data volume of the data of the second data type is less than the data volume of the data of the first data type.

12. The method according to any one of claims 1-11, characterized in that, The method further comprises: The first electronic device receives the screen mirroring content of a second electronic device and the screen mirroring content of a third electronic device; Play the screen mirroring content of the second electronic device at a first bitrate and play the screen mirroring content of the third electronic device at a second bitrate, the first bitrate being less than the second bitrate, and the sum of the first bitrate and the second bitrate being equal to the total bitrate of the first electronic device; wherein the screen mirroring content of the second electronic device is static content, and the screen mirroring content of the third electronic device is dynamic content; or the screen mirroring content of the second electronic device is low-bitrate content, and the screen mirroring content of the third electronic device is dynamic content; or the screen mirroring content of the second electronic device is low-bitrate content, and the screen mirroring content of the third electronic device is static content.

13. The method according to claim 12, characterized in that, The method further comprises: Detect that the screen focus is on the screen mirroring content of the first electronic device; Play the screen mirroring content of the second electronic device at a third bitrate and play the screen mirroring content of the third electronic device at a fourth bitrate, the third bitrate being greater than the fourth bitrate, and the sum of the third bitrate and the fourth bitrate being equal to the total bitrate of the first electronic device.

14. The method according to any one of claims 1 to 13, characterized in that Encoding the first media data specifically includes: Obtain the data stream of the first media data; Divide the data stream into n segments, and further divide each segment of the data stream into k parts, where n is greater than or equal to 1 and k is greater than or equal to 2. The length of the data stream is the time required for encoding the data stream. Perform parallel encoding on the n segments of the data stream. Among them, the encoding start times of the first part of the data stream in the n segments of the data stream are the same. In each segment of the data stream, the encoding start time of the i-th part of the data stream is earlier than that of the (i + 1)-th part of the data stream by timeset, and timeset is less than or equal to the length of each part of the data stream.

15. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored on the memory. The computer program includes a first operating system. The first electronic device is installed with a first application. The first operating system provides: a first interface, a media data interface, and a second capability. The first application is a third-party application. The first application includes program code for calling the first interface. The first interface is used to call the media data interface. The media data interface encapsulates the second capability. The second capability includes: the ability to perform encoding on audio data, the ability to perform encoding on video data, and the ability to perform encoding on audio-video data. Among them, the ability to perform encoding on audio data includes multiple atomic capabilities of audio encoding. The ability to perform encoding on video data includes multiple atomic capabilities of video encoding. The ability to perform encoding on audio-video data includes multiple atomic capabilities of audio-video encoding. The processor executes the computer program to implement the method according to any one of claims 1-14.

16. A computer-readable storage medium, characterized in that, A computer program is stored thereon. The computer program includes a first operating system. The first electronic device is installed with a first application. The first operating system provides: a first interface, a media data interface, and a second capability. The first application is a third-party application. The first application includes program code for calling the first interface. The first interface is used to call the media data interface. The media data interface encapsulates the second capability. The second capability includes: the ability to perform encoding on audio data, the ability to perform encoding on video data, and the ability to perform encoding on audio-video data. Among them, the ability to perform encoding on audio data includes multiple atomic capabilities of audio encoding. The ability to perform encoding on video data includes multiple atomic capabilities of video encoding. The ability to perform encoding on audio-video data includes multiple atomic capabilities of audio-video encoding. When the computer program is executed by the processor, it implements the method according to any one of claims 1-14.

17. A computer program product, characterized in that, The computer program product includes a computer program, the computer program includes a first operating system, a first electronic device is installed with a first application, and the first operating system provides: a first interface, a media data interface, and a second capability; the first application is a third-party application, the first application includes program code for calling the first interface, the first interface is used to call the media data interface, the media data interface encapsulates the second capability, and the second capability includes: the capability of performing encoding on audio data, the capability of performing encoding on video data, the capability of performing encoding on audio and video data, wherein the capability of performing encoding on audio data includes multiple atomized capabilities of audio encoding, the capability of performing encoding on video data includes multiple atomized capabilities of video encoding, and the capability of performing encoding on audio and video data includes multiple atomized capabilities of audio and video encoding; When the computer program is executed by a processor, it implements the method according to any one of claims 1-14.