Electronic device for displaying video synchronized with audio processed by digital signal processor, and method thereof
By utilizing a DSP to decode audio frames and adjust video playback times, the device addresses synchronization challenges, ensuring accurate audio-visual synchronization.
Patent Information
- Application Number
- PCT/KR2025/005742
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-01
- Filing Date
- 2025-04-28
- Publication Date
- 2025-12-26
AI Technical Summary
Existing electronic devices face challenges in synchronizing video and audio playback due to discrepancies between expected and actual playback times, leading to mismatches that affect lip sync performance.
The electronic device employs a digital signal processor (DSP) to decode audio frames and a central processing unit (CPU) to decode video frames, with the DSP estimating playback time discrepancies and communicating these to the CPU to adjust video playback times accordingly, ensuring synchronized output.
This approach effectively compensates for time differences between audio and video playback, achieving precise synchronization and reducing lip sync issues.
Smart Images

Figure KR2025005742_26122025_PF_FP_ABST
Abstract
Description
Electronic device and method for displaying video synchronized with audio processed by a digital signal processor
[0001] The present disclosure relates to an electronic device and method for displaying video synchronized with audio processed by a digital signal processor (DSP).
[0002] The shape and / or size of electronic devices are diversifying. To enhance mobility, electronic devices with reduced size and / or volume are being designed. Electronic devices may include displays and speakers for outputting multimedia content. The multimedia content may include video streamed over a network and / or video recorded by a camera. The multimedia content may also include audio configured to be played along with the video.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0004] In one embodiment, an electronic device may include a display, a speaker, one or more storage media, a memory storing instructions, a first processor including processing circuitry, and a second processor including processing circuitry. The instructions, when executed by the first processor, may cause the electronic device to receive an input for playing multimedia content. The instructions, when executed by the first processor, may cause the electronic device to identify an audio codec associated with audio frames included in the multimedia content, in response to the input. The instructions, when executed by the first processor, may cause the electronic device to obtain a first set of audio frames from the multimedia content. The instructions, when executed by the first processor, may cause the electronic device to identify playback times of the audio frames included in the first set. The instructions, when executed by the first processor, may cause the electronic device to combine time differences between each of the playback times with respect to an expected playback time of an audio frame indicated by the audio codec to obtain a first point in time having a time difference exceeding a specified reference value. The instructions, when executed by the first processor, may cause the electronic device to transmit the first set of audio frames to the second processor. The instructions, when executed by the first processor, may cause the electronic device to control the display to play a video within the multimedia content included in a time interval corresponding to the first set.The instructions, when executed by the second processor, may cause the electronic device to perform decoding of the audio frames based on receiving the first set of audio frames. The instructions, when executed by the second processor, may cause the electronic device to transmit, to the first processor, a second point in time of the PCM frames to be played back through the speaker based on controlling the speaker using PCM frames obtained based on the decoding and each corresponding to the audio frames. The instructions, when executed by the first processor, may cause the electronic device to control playback of the video through the display to at least partially compensate for the time difference based on receiving, from the second processor, a second point in time after the first point in time.
[0005] In one embodiment, a method of an electronic device may be provided. The electronic device may include a display, a speaker, a first processor, and a second processor. The method may include an operation of controlling the first processor to receive an input for playing the multimedia content. The method may include an operation of controlling the first processor to, in response to the input, identify an audio codec associated with audio frames included in the multimedia content. The method may include an operation of controlling the first processor to obtain a first set of the audio frames from the multimedia content. The method may include an operation of controlling the first processor to identify playback times of the audio frames included in the first set. The method may include an operation of controlling the first processor to combine time differences of each of the playback times with respect to an expected playback time of one audio frame indicated by the audio codec, and to obtain a first point in time having a time difference exceeding a specified reference value. The method may include controlling the first processor to transmit the first set of audio frames to the second processor. The method may include controlling the display to play a video within the multimedia content included in a time interval corresponding to the first set, by controlling the first processor. The method may include controlling the second processor to perform decoding on the audio frames based on receiving the first set of audio frames.The method may include an operation of controlling the second processor to control the speaker using pulse code modulation (PCM) frames, each of which is obtained based on the decoding and corresponding to the audio frames, and transmitting a second point in time of the PCM frames to be played back through the speaker to the first processor. The method may include an operation of controlling the first processor to receive a second point in time after the first point in time from the second processor, and controlling the playback of the video through the display to at least partially compensate for the time difference.
[0006] In one embodiment, a non-transitory computer-readable storage medium comprising instructions may be provided. The instructions, when executed by an electronic device including a display, a speaker, a first processor, and a second processor, may cause the electronic device to control the first processor to detect an input for playing multimedia content. The instructions, when executed by the electronic device, may cause the electronic device to check an audio codec of the multimedia content in response to the input. The instructions, when executed by the electronic device, may cause the electronic device to identify a difference between an expected playback time of audio frames included in the multimedia content, indicated by the checked audio codec, and a time at which an audio signal indicated by the audio frames is output from the speaker based on decoding of the audio frames based on the second processor. The instructions, when executed by the electronic device, may cause the electronic device to change the time of the video displayed on the display based on identifying the difference while controlling the display to play the video of the multimedia content.
[0007] In one embodiment, an electronic device may include a first processor including a display, a speaker, processing circuitry, and a second processor including processing circuitry. The second processor may be configured to decode a set of audio frames based on receiving the set of audio frames from the first processor. The second processor may be configured to control a speaker such that an acoustic signal represented by the sequence of pulse coded modulation (PCM) frames is output through the speaker based on identifying a sequence of PCM frames corresponding to each of the audio frames based on the decoding. The second processor may be configured to transmit, to the first processor, information indicating one PCM frame, among the PCM frames, corresponding to the acoustic signal being reproduced through the speaker while controlling the speaker based at least on the sequence of PCM frames. The above information may be used to control playback of video frames being output through the display controlled by the first processor based on a difference between an expected playback time obtained by decoding the audio frames and a playback time indicated by the information.
[0008] Figure 1 illustrates an exemplary operation of an electronic device for synchronizing video and audio of multimedia content.
[0009] FIG. 2 is a schematic diagram of hardware included in an electronic device and software running on the hardware, according to one embodiment.
[0010] Figure 3 illustrates the operation of an application processor (AP) of an electronic device that performs audio offload.
[0011] FIG. 4 illustrates the operation of an electronic device according to one embodiment.
[0012] Figure 5 illustrates pulse coded modulation (PCM) frames decoded from audio frames by an electronic device.
[0013] Figure 6 illustrates the operation of an AP of an electronic device that identifies the time of sound output from a speaker based on audio offload.
[0014] Figure 7 illustrates the operation of an electronic device that adjusts the playback time of a video based at least on the duration of the sound output from the speaker.
[0015] FIG. 8 is a block diagram of an electronic device within a network environment according to various embodiments.
[0016] Figure 9 is a block diagram of an audio module according to various embodiments.
[0017] Hereinafter, various embodiments of this document are described with reference to the attached drawings.
[0018] FIG. 1 illustrates an exemplary operation of an electronic device (101) for synchronizing video and audio of multimedia content. Referring to FIG. 1, the electronic device (101) may include a display (110) and / or a speaker. An exemplary hardware configuration of the electronic device (101) including a display (110), a speaker, and circuitry for controlling them is described below with reference to FIG. 2.
[0019] In one embodiment, the electronic device (101) can generate or output (or play back or display) multimedia content. The multimedia content can be described as video, audio, text, or any combination thereof. The present disclosure describes a combination of video (e.g., a collection of sequentially displayed or sequentially alternated images, referred to as movies, motion pictures, shorts, and / or clips) and audio (e.g., mono, stereo, or multi-channel audio configured to play back in synchronization with video) as an example of multimedia content, but the types and / or categories of multimedia content that can be output by the electronic device (101) are not limited thereto.
[0020] Referring to FIG. 1, an exemplary state of an electronic device (101) outputting multimedia content is illustrated. The multimedia content may be distributed, transmitted, or stored in the form of a file (120). The multimedia content may be distributed (e.g., streaming) in the form of a bitstream via a network (e.g., the Internet). Although the present disclosure describes an exemplary operation of an electronic device (101) outputting multimedia content stored in a file (120), the electronic device (101) may also perform an operation identical to or similar to the operation of outputting multimedia content stored in a file (120) for multimedia content included in a bitstream transmitted to the electronic device (101) via a network in order to output the multimedia content.
[0021] According to one embodiment, the electronic device (101) may detect or receive an input for playing multimedia content. The input may include a gesture for selecting a file (120) in which multimedia content is stored (e.g., a tap gesture, a mouse click, and / or a mouse double-click on an icon representing the file (120). The input may include a combination of a keyword that triggers voice recognition (e.g., “Hey Bixby”) and a utterance that specifies the file (120) in which multimedia content is stored (e.g., “Play (file name)”). The input may include a gaze directed toward the file (120) in which multimedia content is stored, and / or a gesture performed simultaneously with the gaze (e.g., a pinch gesture). The type and / or category of input that the electronic device (101) can receive may be related to the form factor of the electronic device (101), which will be described later with reference to FIG. 2.
[0022] In one embodiment, an electronic device (101) that receives an input for playing multimedia content may identify a file (120) corresponding to the multimedia content specified by the input. The act of identifying the file (120) may include an act of at least partially loading the file (120) into a volatile memory, such as a random access memory (RAM) (e.g., caching and / or paging). The file (120) may include metadata (130), which includes information about the multimedia content, information describing the multimedia content (e.g., information required for outputting the multimedia content and / or a summary of the multimedia content), video data of the multimedia content (e.g., video frames (140)), and audio data of the multimedia content (e.g., audio frames (150)). The electronic device (101) can identify or verify, using metadata (130), a plurality of bits corresponding to video frames (140) (or a portion of the file (120)) and / or a plurality of bits corresponding to audio frames (150) (or another portion of the file (120)) among the bits included in the file (120).
[0023] For example, a file (120) may include video frames (140) and / or audio frames (150) to which a compression technology of video and / or audio, referred to as a codec, is applied. A codec may be described as a combination of an encoder (or coder) that generates compressed data having a smaller size than the raw data from raw data representing video and / or audio, and a decoder that generates the raw data from the compressed data, the raw data being outputtable through a display (110) and / or a speaker. The encoder may be described as hardware, software, or a combination thereof for generating compressed data from raw data. The decoder may be described as hardware, software, or a combination thereof for generating raw data from compressed data.
[0024] For example, the file (120) may include video frames (140) and / or audio frames (150) in a format of compressed data readable by a decoder. For video and audio, respectively, various codecs have been standardized and / or implemented. The electronic device (101) may determine or identify the codecs (or encoders) used to generate (or encode or compress) the video frames (140) and audio frames (150), respectively, from the metadata (130) of the file (120). In the present disclosure, an audio codec may include an encoder and / or decoder required for generating, storing, and / or reproducing audio frames (150). In the present disclosure, a video codec may include an encoder and / or decoder required for generating, storing, and / or reproducing video frames (140).
[0025] In one embodiment, the electronic device (101) may, in response to an input for playing a file (120) representing multimedia content, check or identify codecs (e.g., a video codec and / or an audio codec) of video frames (140) and / or audio frames (150) stored in the file (120). Based on the identified video codec, the electronic device (101) may process the video frames (140). Based on the processing of the video frames (140), the electronic device (101) may obtain raw data of the video, which may be displayed through the display (110). By controlling the display (110) using the raw data, the electronic device (101) may play back or display the video represented by the video frames (140).
[0026] According to one embodiment, the electronic device (101) may process (e.g., parse and / or decode) audio frames (150) of a file (120) based on identifying an audio codec of the audio frames (150) stored in the file (120). The audio frames (150) may include digital information representing an acoustic signal (or vibration) to be output through a speaker in a time interval defined by the audio codec (e.g., a time interval having a duration of several milliseconds (hereinafter, 'ms') to several tens of ms). The audio frames (150) and / or the digital information stored in each of the audio frames (150) may be referred to as encoded information and / or encoded data in terms of being compressed by the audio codec.
[0027] Referring to FIG. 1, five audio frames (e.g., a first audio frame (150-1), a second audio frame (150-2), a third audio frame (150-3), a fourth audio frame (150-4), and / or a fifth audio frame (150-5)) corresponding to respective time segments of audio of multimedia content are exemplarily illustrated. The playback time (or duration) of each audio frame may be defined by an audio codec. For example, if the playback time of an audio frame is defined as 21.333 ms, the five audio frames may include compressed data representing audio for a time segment of 106.667 ms (= 21.333 ms × 5). For example, the electronic device (101) may calculate or estimate the expected playback time of the audio frames (150) based on the audio codec identified from the metadata (130) (or the unit playback time of the audio frame defined in the metadata (130)) and the number and / or size of the audio frames (150).
[0028] In one embodiment, the electronic device (101) may control a speaker based on audio frames (150) (e.g., decoding audio frames (150)) while controlling a display (110) based on video frames (140) (e.g., decoding video frames (140)). Decoding of the video frames (140) and decoding of the audio frames (150) may be performed by different processors (or circuits) within the electronic device (101), respectively. For example, decoding of the audio frames (150) may be performed by a central processing unit (CPU) for general-purpose calculation and / or operation, and a digital signal processor (DSP) (or an audio signal processor (ASP)) included in the electronic device (101). A DSP is a processor (or circuit) dedicated to decoding audio frames (150), and can perform decoding of audio frames (150) with less power and / or at a faster speed than a CPU. For example, video frames (140) included in a file (120) can be processed by a CPU, and audio frames (160) can be processed by a DSP.
[0029] In one embodiment, a method for cooperatively controlling the processors (e.g., CPU and / or DSP) may be required for synchronous output of video and audio included in multimedia content. For example, a method may be required for synchronizing the operation of a first processor (e.g., CPU) for decoding video frames (140) and the operation of a second processor (e.g., DSP) for decoding audio frames (150). For the synchronization, a method for predicting a mismatch between audio output from a speaker of the electronic device (101) and video displayed through a display (110) may be required. The operation of the processors configured to synchronously output video and audio is described with reference to FIGS. 3 and 4.
[0030] In one embodiment, a mismatch of video and audio may be caused by a difference between an expected playback time of audio frames (150), estimated based on metadata (130), and an actual playback time of audio represented by raw data decoded from the audio frames (150). For example, if a first processor, such as a CPU, does not obtain the expected playback time from the audio frames (150) based on demultiplexing, the first processor may not be able to determine the difference between the actual playback time of audio (e.g., PCM data included in a PCM frame) decoded from the audio frames (150) by a second processor, such as a DSP, and the playback time represented by the time information (e.g., timestamp) of the audio frames (150). In the example, the first processor may not be able to identify and reduce the mismatch of video and audio caused by the difference as the audio is output.
[0031] Referring to the exemplary case of FIG. 1, if the playback time of an audio frame is defined as 21.333 ms by metadata (130), the expected playback time of five audio frames (e.g., the first audio frame (150-1) to the fifth audio frame (150-5)) can be estimated to be 106.667 ms (= 21.333 ms χ 5). In the exemplary case, if each of the first audio frame (150-1) to the fifth audio frame (150-5) includes only compressed data for 20 ms of audio, the actual playback time of the entire raw data corresponding to the first audio frame (150-1) to the fifth audio frame (150-5) can be 100 ms (= 20 ms χ 5). That is, the expected playback time and the actual playback time may have a difference of 6.667 ms (= 106.667 ms - 100 ms). This difference may cause a mismatch between the video and audio played on the display (110). An exemplary operation of the electronic device (101) decoding audio frames (150) is described with reference to FIG. 5.
[0032] According to one embodiment, the electronic device (101) may identify or calculate a difference (td) (e.g., 6.667 ms in the exemplary case) between the time (e.g., 100 ms in the exemplary case) at which an acoustic signal represented by the audio frames (150) is output from a speaker, based on an expected playback time (e.g., 106.667 ms in the exemplary case) of audio frames (150) included in the multimedia content, indicated by an audio codec, checked using metadata (130), and decoding of the audio frames (150) (e.g., decoding of the audio frames (150) based on a second processor, different from the first processor that decodes the video frames (140). The electronic device (101), while controlling the display (110) to play a video of multimedia content, can change the time of the video displayed on the display (110) based on identifying the difference. An exemplary operation in which a first processor communicates with a second processor to identify the difference (td) is described with reference to FIG. 6. An exemplary operation in which the first processor changes the time of the video displayed on the display (110) is described with reference to FIG. 8.
[0033] As described above, an electronic device (101) having a processor (e.g., a DSP) for decoding and / or rendering (e.g., outputting an audio signal through a speaker) audio frames (150) with relatively little power (or current) may decode video frames (140) using another processor (e.g., a CPU). A technique for distributed processing of audio frames (150) paired with video frames (140) using a circuit dedicated to decoding and / or rendering audio frames (150), such as a DSP, may be referred to as offload audio (or offload). The DSP may be referred to as an offload device (e.g., offload circuit) from the perspective of a device that performs offload audio. The electronic device (101) may estimate or calculate an error (e.g., a difference between the expected playback time of the audio frames (150) indicated by the audio codec and the actual playback time of each of the audio frames (150)) occurring in the offload device (e.g., a DSP) when decoding the audio frames (150) based on the offload audio. The electronic device (101) that has identified the difference may control the playback of the video through the display (110) to compensate for the difference. Based on the control of the playback of the video, the electronic device (101) may synchronize the times at which the video frames (140) and the audio frames (160) are played (e.g., lip sync).
[0034] Hereinafter, with reference to FIG. 2, an exemplary hardware configuration included in an electronic device (101) for synchronization of video frames (140) and audio frames (150) of FIG. 1 is described.
[0035] FIG. 2 is a schematic diagram of hardware included in an electronic device (101) and software executed based on the hardware, according to one embodiment. Referring to FIG. 2, the electronic device (101) may be one of various forms of electronic devices, such as a laptop PC (personal computer) (290), smartphones (291) having various form factors (e.g., a bar-type smartphone (291-1), a foldable-type smartphone (291-2), or a sliderable (or rollable) type smartphone (291-3)), a tablet PC (292), a head-mounted display (HMD) device (293), a watch (294), a cellular phone (not shown), and other similar computing devices (not shown).
[0036] In one embodiment, the electronic device (101) may be referred to as a mobile device, a user equipment (UE) (or user terminal), a multi-function device, a portable communication device, a portable device, or a server. The form factor of the electronic device (101) is not limited to the exemplary form factors illustrated in FIG. 2. For example, the electronic device (101) may be included as an electronic control unit (ECU) in a vehicle (e.g., an electric vehicle (EV)). For example, the electronic device (101) may have a form factor that is wearable by a user, such as an earbud (or wireless earphone) and / or a ring, or may have a form factor that is implantable on a body part of a user. For example, the electronic device (101) may have a form suitable for playing multimedia content.
[0037] Referring to FIG. 2, according to one embodiment, an electronic device (101) may include an application processor (AP) (210) and / or a memory (205). The electronic device (101) may further include a display (110) and / or a speaker (240). The AP (210) may be electrically and / or operatively coupled with the memory (205), the display (110), and / or the speaker (240). Electrical coupling of electronic components may include a state in which a wired signal path (or a connection for wireless communication) for transmitting a signal is established between the electronic components. Operational coupling of electronic components may include a state in which the electronic components are directly coupled (or a state in which the electronic components are indirectly coupled) such that one of the electronic components controls another electronic component.
[0038] Referring to FIG. 2, for convenience of explanation, electrical connections between the display (110), the speaker (240), the memory (205), and the AP (210) are schematically illustrated. The AP (210) may be communicatively coupled to the display (110), the speaker (240), and / or the memory (205) via one or more electronic components (e.g., a bus, and / or a communication bus). A wired interface for transmitting information may be established between the AP (210), the memory (205), the display (110), and the speaker (240).
[0039] The AP (210) of FIG. 2 may include circuits (e.g., processing circuits and / or cores) for performing operations (e.g., arithmetic operations and / or logical operations) on data. Binary codes (e.g., instructions) representing the operations may be input to the AP (210). The AP (210) may include a CPU (220) and / or a DSP (230). A single package including all of the CPU (220) and the DSP (230) may be referred to as a system on a chip (SoC). The CPU (220) may have a structure (e.g., a multi-core structure based on a combination of multiple core circuits such as a dual core, a quad core, a hexa core, or an octa core) for loading (or fetching) and / or executing multiple instructions simultaneously.
[0040] In one embodiment, the CPU (220) may include processing circuitry for performing various functions indicated by the instructions. The DSP (230) may include processing circuitry for digital signal processing (e.g., decoding) of audio frames, as described above with reference to FIG. 1. The functions and / or operations described with reference to the present disclosure may be individually or collectively performed by processors included in the AP (210) (e.g., the CPU (220) and / or the DSP (230)). The processors may execute instructions stored in the memory (205) to perform the functions and / or operations indicated by the instructions.
[0041] Referring to FIG. 1, an embodiment in which the DSP (230) is included in the AP (210) is illustrated, but the location of the DSP (230) is not limited thereto. For example, the DSP (230) may be placed outside the AP (210) (e.g., on the PCB of the electronic device (101)) and may be electrically connected to the AP (210) and / or the CPU (220). For example, the DSP (230) may be separated from the electronic device (101). For example, the DSP (230) may be electrically connected to the electronic device (101) through a wired interface (or wireless network) outside the electronic device (101). Through the wired interface (or wireless network), the DSP (230) may be connected to the AP (210) and / or the CPU (220) within the electronic device (101).
[0042] The memory (205) of FIG. 2 may include a circuit for storing data (or instructions) input to or output from the AP (210). The memory (205) may include a volatile memory such as a random-access memory (RAM), and / or a non-volatile memory such as a read-only memory (ROM). The non-volatile memory may be referred to as storage. The volatile memory may include, for example, at least one of a dynamic RAM (DRAM), a static RAM (SRAM), a cache RAM, and a pseudo SRAM (PSRAM). The non-volatile memory may include, for example, at least one of a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a flash memory, a hard disk, a compact disc, a solid state drive (SSD), and an embedded multimedia card (eMMC). The memory (205) may include one or more storage media (e.g., the volatile memory and / or non-volatile memory described above) distributedly located in the electronic device (101). The AP (210) of the electronic device (101) may execute instructions of the memory (205) within the electronic device (101) to perform functions and / or operations (e.g., the operations of FIG. 4) indicated by the instructions. For example, circuits included in the AP (210) (e.g., the CPU (220) and / or the DSP (230)) may be configured to collectively or individually execute the instructions.
[0043] The display (110) of the electronic device (101) may include a circuit for visualizing information provided from the AP (210). The display (110) may include a liquid crystal display (LCD), a plasma display panel (PDP), and / or light emitting diodes (LEDs). The LEDs may include organic LEDs (OLEDs). The embodiment is not limited thereto, and the display (110) may include electronic paper. The display area (or active area) of the display (110) may include an area where light is emitted, formed by pixels (e.g., activated pixels) of the display (110). The display (110) may include a sensor (e.g., a touch sensor) for detecting an external object (e.g., a user's finger) on the display (110). The sensor may be included in the display (110) in the form of a panel (e.g., a touch sensor panel (TSP)).
[0044] Referring to FIG. 2, programs executed by the CPU (220) (e.g., a playback unit (251), a demultiplexer (252), a video output unit (253), an offload data transmitter (254), a media clock unit (255), a playback time management unit (257), and / or an error calculation unit (256)) are illustrated. The programs may be independently installed in the memory (205), or may be stored in the memory (205) as sub-routines (or applets or dynamic link libraries (DLLs)) of a single program. In the present disclosure, the remaining 'units', excluding the 'units' of the playback unit (251), the video output unit (253), the offload data transmitter (254), the media clock unit (255), the playback time management unit (257), and the error calculation unit (256), may be used to refer to hardware included in the electronic device (101).
[0045] Referring to FIG. 2, the playback unit (251) may be a program for detecting or processing events (e.g., playback, pause, stop, fast forward, and / or rewind of multimedia content) occurring in a software application (e.g., a gallery application and / or a video streaming application) related to multimedia content. Based on the execution of the playback unit (251), the AP (210) (or CPU (220)) may manage other programs (e.g., a demultiplexer (252), a video output unit (253), an offload data transmitter (254), a media clock unit (255), a playback time management unit (257), and / or an error calculation unit (256)) illustrated in FIG. 2.
[0046] Referring to FIG. 2, a demultiplexer (or demuxer) (252) may be a program for performing demultiplexing on a media content source. In one embodiment, the demultiplexing performed by executing the demultiplexer (252) may include an operation of parsing, extracting, and / or obtaining one or more audio frames (e.g., audio frames (150) of FIG. 1) from the multimedia content. The parsing of the multimedia content performed by the demultiplexer (252) may include an operation of extracting video frames (e.g., video frames (140) of FIG. 1) and / or audio frames (e.g., audio frames (150) of FIG. 1) from the multimedia content. The media content source may include information stored in the memory (205), such as a file (120) of FIG. 1. The file (120) may be configured to store multimedia content based on the MPEG (motion picture expert group)-4 format and / or the MKV (matroska multimedia container) format. The media content source may include a bitstream streaming from the Internet, a camera, and / or a microphone. Although not shown, the electronic device (101) may further include a communication circuit to receive a bitstream from the Internet.
[0047] In one embodiment, the AP (210) (or CPU (220)) executing the demultiplexer (252) may store (e.g., buffer and / or cache) audio frames extracted from multimedia content in a buffer (e.g., audio buffer) of the memory (205). The buffer may be allocated by the AP (210) executing the demultiplexer (252). Based on the limited size of the buffer, the AP (210) may extract audio frames from the multimedia content. For example, the AP (210) may obtain or load a set of audio frames having the size of the buffer from the multimedia content.
[0048] Referring to FIG. 2, the video output unit (253) may be a program that performs decoding on video frames (or video streams) parsed by the demultiplexer (252). Although the operation of the CPU (220) performing decoding on video frames is exemplarily described, the embodiment is not limited thereto, and a circuit dedicated to performing decoding on video frames may be included in the AP (210). The AP (210) (or CPU (220)) executing the video output unit (253) may control the display (110) to output raw data (e.g., video frames in raw format) decoded from the video frames. For example, the AP (210) may control the output of video frames through the display (110) according to a refresh rate (e.g., frame rate) set by the multimedia content (or set by the video codec) based on a media clock. The above playback rate can be expressed as a numeric value with units of fps (frames per second) and / or Hz.
[0049] Referring to FIG. 2, the media clock unit (255) may be a program that calculates the time of an audio signal being played through a speaker (240) controlled by the DSP (230). For example, the AP (210) (or CPU (220)) executing the media clock unit (255) may calculate or obtain the time (e.g., output time) of an audio signal currently being played through the speaker (240) by combining the time stamp of the initial audio frame transmitted to the DSP (230) for offloading with the time of the audio signal output after the initial audio frame.
[0050] Referring to FIG. 2, the offload data transmitter (254) may be a program for transmitting at least one audio frame (or a set of audio frames) parsed by the demultiplexer (252) to an offload device, such as a DSP (230). The AP (210) (or CPU (220)) executing the offload data transmitter (254) may transmit a set of audio frames stored in a buffer allocated to the memory (205) (e.g., stored in a size equal to the size of the buffer) to the DSP (230). For example, if a buffer having a size of 256 kilobytes (KB) is allocated in the memory (205), the CPU (220) may store a number of audio frames whose total size is 256 KB or less in the buffer of the memory (205). For example, if the size of an audio frame is 512 bytes, the CPU (220) can store 500 (= 256 KB / 512 bytes) audio frames in the buffer and transmit the 500 audio frames to the DSP (230).
[0051] In one embodiment, while multimedia content is being played, the offload data transmitter (254) may repeatedly (or periodically) transmit the audio frames stored in the buffer to the DSP (230) when the number of audio frames corresponding to the size of the buffer is stored in the buffer. After all audio frames stored in the buffer are transmitted to the DSP (230), the AP (210) executing the offload data transmitter (254) may remove the audio frames stored in the buffer (e.g., reset and / or flush the buffer). Transmitting the audio frames from the buffer to the DSP (230) may be performed based on a function for driving the CPU (220) independently of the transmission of information, such as direct memory access (DMA).
[0052] In one embodiment, the offload data transmitter (254) may transmit a set of audio frames loaded by the demultiplexer (252) to an offload device (e.g., a DSP (230)). For example, instead of transmitting only one audio frame to the offload device, audio frames filled in a buffer having a specific size (e.g., the size of the buffer) may be transmitted in batches. For example, audio frames stored in a buffer having a size of 256 KB may be copied to and / or transmitted to the DSP (230) by the CPU (220).
[0053] In one embodiment, in response to audio frames transmitted by the offload data transmitter (254), the DSP (230) may decode the audio frames. The DSP (230) may perform decoding of the audio frames according to the order in which each of the audio frames is to be played. The decoded audio frame may be referred to as raw data for the audio frame, and / or a PCM frame. The PCM frame may be described as uncompressed linear audio data. The DSP (230) may control the speaker (240) based on the PCM frame, thereby causing the speaker (240) to output an acoustic signal having vibration and / or intensity represented by the PCM frame. Rendering based on the DSP (230) may include outputting an acoustic signal represented by the decoded raw data (e.g., the PCM frame) through the speaker (240).
[0054] Referring to FIG. 2, the error calculation unit (256) may be a program that can (cumulatively) calculate or measure the difference (e.g., error) between the expected playback time and the actual playback time of an audio frame, which occurs while decoding and / or rendering based on the DSP (230) is performed. The AP (210) (or CPU (220)) that executes the error calculation unit (256) may obtain or identify one or more parameters (e.g., a timestamp and / or a playback time of the audio frame) required to calculate the expected playback time of the audio frame parsed by the demultiplexer (252) by executing the demultiplexer (252). The AP (210) that executes the error calculation unit (256) may identify the expected playback time of the audio frame by using an audio codec to be used for decoding the audio frame and / or a profile of the audio codec.
[0055] In one embodiment, the AP (210) executing the error calculation unit (256) can determine the playback time of an audio signal derived from an audio frame from the metadata of the audio frame (or the header of the audio frame). The audio signal may be an audio signal represented by a PCM frame that is decoded by the DSP (230) and corresponds to the audio frame. The AP (210) can determine whether the playback time (e.g., actual playback time) determined for the audio signal matches the expected playback time of the audio frame. If the playback time determined from the audio signal does not match the expected playback time, the AP (210) can obtain the difference between the playback time and the expected playback time. The difference may cause a mismatch between the audio frame and the video frame corresponding to the audio frame when rendering the audio frame.
[0056] In one embodiment, the AP (210) can determine when to synchronize the video frame and the audio frame, and the degree to which to lag or lead the video frame at the time when the difference between the expected playback time of the audio frame, accumulated by the error calculation unit (256), and the actual playback time increases beyond a specified reference value.
[0057] In one embodiment, the playback time management unit (257) may be a program that adjusts parameters for controlling playback of video frames, such as a media clock, to synchronize video frames and audio frames. The AP (210) (or CPU (220)) executing the playback time management unit (257) may check the expected playback time. The AP (210) executing the playback time management unit (257) may identify or load the time at which to synchronize the video frames and audio frames, calculated by the error calculation unit (256). When the loaded time is reached, the AP (210) executing the playback time management unit (257) may check the time of the audio signal being output from the offload device. The AP (210) may adjust the time of the media clock used for playback of the video frames based on the time of the audio signal. Based on the adjustment of the time of the media clock, the AP (210) may control playback of the video frames.
[0058] Below, with reference to FIG. 3, information transmitted between the CPU (220) and the DSP (230) in a state of outputting multimedia content is described.
[0059] FIG. 3 illustrates the operation of an application processor (AP) (e.g., AP (210) of FIG. 2) of an electronic device that performs audio offload. The electronic device of FIG. 3 may include the electronic device (101) of FIG. 1 and / or FIG. 2. The CPU (220), DSP (230), and speaker (240) of FIG. 3 may correspond to the CPU (220), DSP (230), and speaker (240) of FIG. 2, respectively.
[0060] Referring to FIG. 3, the CPU (220) can execute a software application (310). The software application (310) can provide a function for playing multimedia content. By executing the software application (310), the CPU (220) can display a user interface (UI) capable of receiving an input for playing multimedia content. The UI can be displayed on a display (e.g., the display (110) of FIG. 1 and / or FIG. 2). Through the UI, the CPU (220) can receive a request and / or input for playing multimedia content.
[0061] For example, a CPU (220) that receives a request to play multimedia content can identify the multimedia content by executing a player framework (320). The player framework (320) can include programs (e.g., a playback unit (251), a demultiplexer (252), a video output unit (253), an offload data transmission unit (254), a media clock unit (255), an error calculation unit (256), and a playback time management unit (257)) illustrated in FIG. 2. The CPU (220) that executes the player framework (320) can obtain a set of audio frames included in the multimedia content. The set can be stored in a buffer allocated to a memory (e.g., a memory (205) of FIG. 2).
[0062] Referring to FIG. 2, a set of audio frames acquired based on the player framework (320) can be processed by a CPU (220) executing a kernel (330). The CPU (220) executing the kernel (330) can transmit the set of audio frames to the DSP (230). The CPU (220) can transmit the set of audio frames to the DSP (230) and may not transmit information indicating the time at which the audio frames are output (e.g., timestamp information).
[0063] Referring to FIG. 3, when the DSP (230) receives a set of audio frames from the CPU (220), it can perform decoding and / or rendering on the set of audio frames. For example, the DSP (230) can obtain raw data (e.g., PCM frame) corresponding to the audio frame and transmit the raw data to the speaker (240). For example, the DSP (230) can transmit the raw data to the speaker (240) so that an acoustic signal represented by the raw data is output through the speaker (240).
[0064] In one embodiment, the CPU (220) may request the DSP (230) for time information about an audio signal being output by the DSP (230) in order to perform synchronization between video frames and audio frames based on an expected playback time. The request may be performed based on the execution of the player framework (320). In response to the request, the DSP (230) may transmit or return to the CPU (220) the time corresponding to the currently output audio signal. For example, the DSP (230) may transmit to the CPU (220) time information (e.g., duration) corresponding to a PCM frame associated with an audio signal being output through a speaker (240). The time information transmitted from the DSP (230) may be related to the size of audio frames decoded by the DSP (230).
[0065] In one embodiment, the CPU (220) may obtain information about an audio frame rendered by the DSP (230) (e.g., a timestamp indicating the time of an audio signal output through a speaker (240)) from the DSP (230). The CPU (220) may compare the information with the expected playback times of the audio frames to perform synchronization between the video frames and the audio frames. The synchronization may be performed to reduce mismatch between the video frames and the audio frames generated in the DSP (230) while maintaining low-power decoding and low-power rendering of the audio frames by the DSP (230). Hereinafter, an exemplary flowchart of the CPU (220) performing synchronization between the video frames and the audio frames is described with reference to FIG. 4.
[0066] FIG. 4 illustrates the operation of an electronic device according to one embodiment. The electronic device of FIG. 4 may include the electronic device (101) of FIG. 1 and / or FIG. 2. The operations of FIG. 4 may be performed by the AP (210) and / or the CPU (220) of FIG. 2.
[0067] Referring to FIG. 4, in operation (410), according to one embodiment, an electronic device may receive an input for playing multimedia content. The input may be received via a software application for playing media. The electronic device receiving the input may execute the playback unit (251) of FIG. 2. The electronic device receiving the input may transmit multimedia content corresponding to the input to the playback unit (251) of FIG. 2. For example, the electronic device may transmit a file representing the multimedia content (e.g., file (120) of FIG. 1) and / or a bitstream to the playback unit (251) of FIG. 2.
[0068] Referring to FIG. 4, in operation (415), according to an embodiment, an electronic device may determine to offload audio of multimedia content. A first processor, such as a CPU of the electronic device (e.g., CPU (220) of FIG. 2), may perform operation (415). For example, the electronic device may determine whether to perform offload audio in order to output the multimedia content. The electronic device may obtain information about audio data included in the multimedia content (e.g., a bitstream, referred to as an audio stream, including one or more audio frames) using the demultiplexer (252) of FIG. 2. From the information, the electronic device may identify at least one of a sampling rate of the audio data, a codec used to encode the audio data, a channel of the audio data (e.g., a mono channel, a stereo channel, or three or more multi-channels), and / or a PCM sample size. By comparing the information with conditions defined for offload audio, the electronic device may determine whether to perform offload audio.
[0069] For example, if audio data included in multimedia content is generated based on an audio codec supported by an offload device (e.g., a second processor such as the DSP (230) of FIG. 2), the electronic device may decide to perform offload audio. To determine the offload of operation (415), the electronic device may identify an audio codec corresponding to audio frames included in the multimedia content. For example, in response to an input of operation (410), the electronic device may identify an audio codec associated with audio frames included in the multimedia content. If it is determined to perform offload audio, the electronic device may perform operation (420).
[0070] Referring to FIG. 4, in operation (420), according to one embodiment, an electronic device may determine an expected playback time using a profile of an audio codec associated with multimedia content. The profile may be identified or parsed based on the demultiplexer (252) of FIG. 2. The electronic device may determine the expected playback time based on a PCM sample size per audio frame defined for an audio codec indicated by the profile. The profile of operation (420) may be included in metadata (e.g., metadata (130) of FIG. 1) of a file (e.g., file (120) of FIG. 1) including multimedia content. The expected playback time of operation (420) may be related to the size of PCM samples included in an audio frame of the multimedia content, which is indicated by the profile of the audio codec. For example, depending on the audio codec, the PCM sample size per audio frame may be defined as in Table 1.
[0071] Audio codec name PCM sample size per audio frame AAC (Advanced Audio Coding) LC (Low Complexity) 1024 AAC SBR (Spectral Band Replication) (HE (high efficiency)-AAC) 2048 AAC PS (parametric stereo) (HE-AAC v2) 2048 MPEG Layer I 384 MPEG Layer II 1152 MPEG Layer III version 11152 MPEG Layer III version 2, 2.5576
[0072] Using the mapping table and / or information shown in Table 1, the electronic device can identify or verify the PCM sample size corresponding to the audio codec. The information stored in the electronic device is not limited to the audio codecs shown in Table 1.
[0073] The PCM sample size in Table 1 may indicate the size of PCM samples included in a PCM frame obtained by decoding an audio frame of a specific audio codec. An electronic device that identifies the PCM sample size in Table 1 may calculate or determine an expected playback time for one audio frame included in multimedia content based on Mathematical Expression 1.
[0074]
[0075] For example, if an electronic device identifies information indicating that audio frames included in the multimedia content are encoded based on AAC LC from metadata stored in the multimedia content (e.g., metadata (130) of FIG. 1), and the sampling rate is 44100 Hz, the expected playback time calculated based on mathematical expression 1 may be 23219 microseconds (hereinafter, 'μs') (= (1024 Х 1000000) / 44100 ).
[0076] Referring to FIG. 4, in operation (425), according to an embodiment, an electronic device may obtain an audio frame (e.g., audio frames 150 of FIG. 1) based on demultiplexing of multimedia content. For example, the electronic device may obtain a set of audio frames from the multimedia content. The audio frame obtained based on operation (425) may include a playback time of the audio frame (e.g., a duration of the audio frame, represented by a numerical value in units of μs) and / or a timestamp of the audio frame (e.g., a time corresponding to the audio frame). The duration and / or the timestamp may be stored in a header (or header area, metadata) of the audio frame. The timestamp may indicate a position corresponding to the audio frame within the entire time interval of the multimedia content.
[0077] An electronic device that obtains an audio frame based on the action (425) can identify or confirm the playback time of the audio frame before decoding the audio frame (or without decoding the audio frame). For example, the electronic device can identify or display the current playback time (or current playback position) using the timestamp of the audio frame. For example, the electronic device can determine the rendering of a video frame using the timestamp of the audio frame. For example, an electronic device that obtains a timestamp indicating 3000 ms from an audio frame can render a video frame having the 3000 ms timestamp and display a visual object (e.g., a UI element referred to as text and / or a slider) on the display indicating that 3000 ms is the current playback time.
[0078] Referring to FIG. 4, in operation (430), according to one embodiment, the electronic device may determine or confirm whether the expected playback time of operation (420) matches the playback time of the audio frame obtained based on operation (425). For example, the electronic device may identify the playback time(s) of the audio frame(s) obtained based on operation (425). From the header (or metadata) of the audio frame, the electronic device may determine the PCM sample size and sampling rate of the audio frame. From the header of the audio frame, the electronic device may obtain time information of the audio frame. The time information obtained from the header of the audio frame may include a timestamp of the audio frame and / or a timestamp difference between audio frames (e.g., a difference value between a timestamp of a current audio frame and a timestamp of a previous audio frame). Using the time information, the electronic device may calculate or identify the playback time of the audio frame. Based on the PCM sample size and / or the sampling rate, the electronic device can calculate the playback time of the audio frame. The playback time can be obtained based on mathematical expression 1.
[0079] For example, the expected playback time obtained based on information included in the profile of the audio codec and mathematical expression 1 may be different from the playback time (e.g., actual playback time) obtained based on information included in the header (or metadata) of the audio frame and mathematical expression 1. For example, if the PCM sample size identified based on the name (or type) of the audio codec and Table 1 (or a mapping table representing Table 1) is different from the PCM sample size indicated by the header of the audio frame, the expected playback time and the actual playback time may be different from each other.
[0080] Referring to FIG. 4, if the expected playback time and the playback time of the audio frame are the same (430 - Yes), the electronic device can continue to perform operation (425) without operations (435, 440, 445). If the expected playback time and the playback time of the audio frame are different (430 - No), the electronic device can perform operation (435).
[0081] Referring to FIG. 4, in operation (435), according to one embodiment, an electronic device may accumulate a time difference between an expected playback time and a playback time (e.g., an audio frame duration). The electronic device may allocate a variable for storing the time difference to a memory (e.g., a volatile memory and / or a register of the electronic device, including the memory (205) of FIG. 2). The electronic device may store a numerical value representing the time difference in the variable. If the expected playback time is calculated as 23219 μs and the playback time of the audio frame is calculated as 23651 μs, the electronic device may store the difference between the expected playback time and the playback time, which is 432 μs (= 23651 μs - 23219 μs), in the variable. The variable may have a name such as AccumulatedDuration. The embodiment is not limited thereto.
[0082] In one embodiment, when the electronic device identifies a time difference between the expected playback time and the playback time of the operation (430), the electronic device may combine the identified time difference with a numerical value that has been accumulated since before the identification. The numerical value may be initialized at the time of starting playback of the multimedia content. The numerical value may be initialized or changed when performing a position movement (e.g., seek) of the multimedia content (or video or audio). For example, the electronic device that receives an input for changing the time of the multimedia content currently being displayed through the electronic device may initialize the numerical value. A variable set to store the numerical value may be created in memory at the time or declared. For example, the numerical value stored in the variable may be defined as in Mathematical Expression 2.
[0083]
[0084] Referring to FIG. 4, in operation (440), according to one embodiment, the electronic device may determine whether the accumulated time difference based on operation (435) exceeds a threshold value. If the accumulated time difference based on operation (435) is less than or equal to the threshold value of operation (440) (440-No), the electronic device may perform operation (425). If the accumulated time difference based on operation (435) exceeds the threshold value of operation (440) (440-Yes), the electronic device may perform operation (445). The threshold value of operation (440) may be determined heuristically based on whether a user can recognize a mismatch between a video frame and an audio frame. For example, according to ITU R BT.1359-1, a user can identify a mismatch when there is a time difference of -100 ms to +40 ms. In the above example, the threshold value may be determined as 40 ms. The embodiment is not limited thereto, and the reference value may be determined to be 90 ms, combined with the error range.
[0085] Referring to FIG. 4, based on operations (425, 430, 435, and 440), each time audio frames are acquired, the difference between the expected playback time and the playback time may be accumulated. For example, the time differences of each of the playback times of the audio frames relative to the expected playback time may be combined within a variable set to accumulate the difference. If the combination of the time differences exceeds the threshold value of operation (440), the electronic device may perform operation (445).
[0086] Referring to FIG. 4, in operation (445), according to one embodiment, the electronic device may store a timestamp of an audio frame that causes an accumulated time difference to exceed a reference value. For example, the electronic device may obtain a first point in time having a time difference exceeding the reference value of operation (440) by combining the time differences of each of the playback times with respect to the expected playback time of one audio frame indicated by the audio codec (e.g., the expected playback time of operation (420)). The timestamp of operation (445) may indicate the first point in time.
[0087] In one embodiment, the electronic device may determine or store a timestamp of operation (445) and a value (e.g., a time adjustment value) to be used to control playback of video frames at the timestamp. For example, a pair of the timestamp and the time adjustment value may be stored in a memory of the electronic device.
[0088] In one embodiment, the time adjustment value may have a value that is difficult for a user to perceive because it causes unnatural playback (e.g., fast-forwarding and / or rewinding) of video frames. For example, the time adjustment value may be defined as a fixed value (e.g., 30000 μs). For example, the time adjustment value may be determined as the accumulated time difference if the accumulated time difference does not exceed a specified threshold, and may be determined as the specified threshold if the accumulated time difference exceeds the specified threshold. In one embodiment, if the time adjustment value is different from the accumulated time difference based on operation (435), the numerical value stored in the variable in which the time difference is accumulated may be deducted by the time adjustment value.
[0089] Referring to FIG. 4, obtaining audio frames based on operation (425) may be performed repeatedly based on the size of a buffer allocated in the memory. In operation (450) of FIG. 4, the electronic device according to one embodiment may transmit a set of audio frames obtained based on the repeatedly performed operation (425) to a second processor (e.g., DSP (230) of FIG. 2). For example, the first processor of the electronic device may execute a kernel as described above with reference to FIG. 3 to transmit the set of audio frames to the second processor. The electronic device may control a display (e.g., display (110) of FIG. 1 and / or FIG. 2) to play a video within multimedia content included in a time interval corresponding to the set of audio frames.
[0090] Based on the operation (450), a second processor (e.g., DSP (230) of FIG. 2) that has received a set of audio frames may perform decoding on the audio frames. Based on the decoding, the second processor may obtain PCM frames, each corresponding to the audio frames. The second processor may control a speaker (e.g., speaker (240) of FIG. 2) using the PCM frames to output an acoustic signal represented by the PCM frames through the speaker. Although the operation of the second processor controlling the speaker is described, the embodiment is not limited thereto, and the second processor may transmit (either wired or wirelessly) the PCM frames to an external electronic device (e.g., a speaker, wireless headphones, and / or earbuds) connected via a wired interface, or via a wireless network such as Bluetooth. For example, the second processor may transmit the PCM frames to an external speaker (wired or wirelessly) connected to the electronic device to output audio represented by the PCM frames through the external speaker.
[0091] Referring to FIG. 4, in operation (455), according to one embodiment, a first processor of an electronic device may identify a time of an audio signal being output through a speaker from a second processor. For example, the second processor may transmit or report, to the first processor, a point in time and / or time of a PCM frame being played through the speaker based on a request and / or a specified period of the first processor, as described below with reference to FIG. 6. The time transmitted from the second processor may transmit a numerical value (e.g., a timestamp) representing a difference between a time at which the output of the audio signal started and a current time. The first processor of the electronic device may calculate or determine a point in time (e.g., a position of the audio signal within the total playback time) of the audio signal being output through the speaker by combining the timestamp corresponding to the first audio frame transmitted to the second processor and the timestamp transmitted from the second processor. The identified point in time can be processed by the media clock unit (255) and / or the video output unit (253) of FIG. 2 and used for outputting and / or rendering video frames.
[0092] Referring to FIG. 4, in operation (460), according to one embodiment, the electronic device may determine or confirm whether the time identified based on operation (455) is greater than or equal to the timestamp stored in operation (445). For example, the electronic device may confirm whether the time of the audio signal being output from the second processor corresponds to a time exceeding the timestamp of operation (445). The fact that the time of the audio signal being output from the second processor exceeds the timestamp of operation (445) may indicate that a mismatch exceeding a reference value has occurred. If the time identified based on operation (455) is less than the timestamp of operation (445) (460-No), the electronic device may not perform operation (465) and may perform operation (455). If the time identified based on the action (455) is greater than or equal to the timestamp of the action (445) (460-Yes), the electronic device may perform the action (465). If the timestamp of the action (445) is not stored (e.g., the action (445) is not performed), the electronic device may not perform any of the actions (460, 465).
[0093] Referring to FIG. 4, in operation (465), according to one embodiment, the electronic device may adjust the playback time of a video of multimedia content. If a time after the timestamp of operation (445) is identified based on operation (455), the electronic device may control playback of the video frames through the display to at least partially compensate for the time difference between the video frames and audio frames being output through the display and the speaker, respectively.
[0094] For example, if a pair of a timestamp and a time adjustment value is stored within the operation (445), the electronic device can change the position and / or order of a video frame to be displayed on the display within a sequence of video frames included in the multimedia content, based on the time adjustment value. For example, if a video frame corresponding to 2200000 μs is being output within the entire time interval of the multimedia content, the electronic device can output a video frame at a point in time (e.g., 2170000 μs = 2200000 μs - 30000 μs) from which the time adjustment value (e.g., 30000 μs) has been subtracted. In order to output (or render) the video frame at the point in time from which the time adjustment value has been subtracted, the electronic device can change the media clock (or the time of the media clock) managed by the media clock unit (255). When outputting a video frame at a point in time from which a time adjustment value has been deducted, the electronic device may also update a variable representing the accumulated time difference based on operation (435). For example, the electronic device may deduct a numerical value stored in the variable by the time adjustment value.
[0095] As described above, in one embodiment, the first processor of the electronic device may determine whether to synchronize video frames and audio frames using information transmitted from the second processor (e.g., information indicating the time of operation (455)) when performing an offload. The first processor may predict or calculate, prior to transmitting a set of audio frames to the second processor, a mismatch between the video frames and the audio frames that will occur when the audio frames are played back by the second processor. The calculated mismatch may be compensated for based on a lag or lead of the video frames being displayed on the display. The degree to which the video frames are shifted in the time domain may be determined within a range that is difficult for a user to perceive.
[0096] Hereinafter, with reference to FIG. 5, an exemplary operation of the second processor for decoding audio frames transmitted to the second processor based on operation (450) is described.
[0097] FIG. 5 illustrates pulse coded modulation (PCM) frames (510-1, 510-2, 510-3, 510-4, 510-5) decoded from audio frames (150-1, 150-2, 150-3, 150-4, 150-5) by an electronic device. The electronic device of FIG. 5 may include the electronic device (101) of FIG. 1 and / or FIG. 2. The operations of FIG. 5 may be performed by the AP (210) and / or the DSP (230) of FIG. 2.
[0098] Referring to FIG. 5, audio frames (e.g., a first audio frame (150-1) to a fifth audio frame (150-5)) obtained by an electronic device by performing demultiplexing on multimedia content are illustrated. The electronic device, which confirms information (e.g., metadata (130) of FIG. 1) indicating that the audio frames are encoded based on an AAC-LC profile from the multimedia content, can confirm that the PCM sample size of each of the audio frames is 1024 based on Table 1. When a sampling rate of 44100 Hz is confirmed from the information, the expected playback time of the audio frame indicated by the information can be calculated as 21.333 ms based on Mathematical Expression 1.
[0099] For example, during the generation of multimedia content (e.g., multiplexing), the information may be changed for synchronization of video frames and audio frames, or audio frames may be generated such that the actual playback time of the audio frame is different from the expected playback time indicated by the information. For example, if the PCM sample size in the information is set to 1043 instead of 1024, which is the PCM sample size corresponding to the AAC-LC codec in Table 1, the difference between the expected playback time and the actual playback time may increase as shown in Table 2 as audio frames are decoded. Both the expected playback time and the actual playback time can be calculated based on Equation 1.
[0100] Audio frame number (identified by the demultiplexer) PCM sample size (identified by the demultiplexer) Duration of audio frame (e.g., expected duration) Accumulated timestamps of audio frames Sample size of PCM frame, decoded in DSP Accumulated timestamps in DSP 1104323651010240210432365123651102423220310432365147302102446440410432365170952102469660510432365194603102592880:201043236514493651024441179
[0101] The playback time of the audio frame identified by the demultiplexer of Table 2 can be calculated based on the profile of the first audio frame (150-1) to the fifth audio frame (150-5) of FIG. 5. The sample size of the PCM frame decoded by the DSP of Table 2 can represent the size of the PCM sample included in the first PCM frame (510-1) to the fifth PCM frame (510-5) of FIG. 5.
[0102] A PCM frame obtained by decoding an audio frame may include PCM bytes. The PCM frame size in Table 2 can be determined as in mathematical expression 3.
[0103]
[0104] Channel Count in Equation 3 may be 2 for stereo channels. PCM Bytes in Equation 3 may represent the byte size of a PCM frame obtained by decoding an audio frame. Sample size in Equation 3 may be the size of a PCM sample expressed in bytes (e.g., 3 for 24 bits, 2 for 16 bits, 3 for 8 bits). For example, if Channel Count is 2, the size of a PCM sample is 16 bits, the PCM sample size is 1024, and the sampling rate is 44100 Hz, the actual playback time of the PCM frame may be 23219 μs (= 1024 Х 1000000 / 44100) based on Equation 1. Referring to Table 2, the accumulated timestamps in the DSP can be accumulated for each audio frame as much as 23219 μs (represented as 23220 based on rounding).
[0105] In one embodiment, the DSP may sequentially transmit a first PCM frame (510-1), a second PCM frame (510-2), a third PCM frame (510-3), a fourth PCM frame (510-4), and a fifth PCM frame (510-5) to a speaker (e.g., speaker (240) of FIG. 2) to output an audio signal represented by the sequence of the PCM frames.
[0106] Referring to Table 2, when decoding the 20th audio frame, the difference between the accumulated expected playback time tracked by the CPU (449365 μs) and the accumulated actual playback time played by the DSP (4411790 μs) may increase to more than 8000 μs. The difference may increase while audio frames continue to be decoded (e.g., while multimedia content continues to be played). For example, the mismatch between video frames and audio frames may increase.
[0107] In one embodiment, an electronic device may track a difference between an expected playback time of Table 2 and a timestamp (e.g., an actual playback time) accumulated in a DSP that decodes the audio frames when performing demultiplexing on multimedia content. For example, a CPU of the electronic device may track the difference by performing the operations of FIG. 4. If the difference exceeds a specified threshold, the CPU may control playback of video frames to compensate for the difference. To track the difference, the CPU may obtain a timestamp accumulated by the DSP, as illustrated in Table 2, from the DSP. Hereinafter, an exemplary operation of the CPU obtaining a timestamp from the DSP is described with reference to FIG. 6.
[0108] FIG. 6 illustrates the operation of an AP of an electronic device that identifies the timing of sound output from a speaker based on audio offload. The electronic device of FIG. 6 may include the electronic device (101) of FIG. 1 and / or FIG. 2. The CPU (220) and DSP (230) of FIG. 6 may correspond to the CPU (220) and DSP (230) of FIG. 2, respectively. The software application (310), the player framework (320), and the kernel (330) of FIG. 6 may correspond to the software application (310), the player framework (320), and the kernel (330) of FIG. 3, respectively.
[0109] In one embodiment, the CPU (220) may transmit a set of audio frames to the DSP (230) (e.g., based on the operations described above with reference to FIGS. 3 and / or 4). While the DSP (230) is decoding and / or rendering the audio frames included in the set, the CPU (220) may request the DSP (230) to transmit the time of the audio frame being decoded (or rendered) by the DSP (230). For example, the CPU (220) may transmit a signal to the DSP (230) requesting transmission of the time. Referring to FIG. 6, the CPU (220) may generate the signal based on the execution of the player framework (320). The signal generated by the player framework (320) may be transmitted to the DSP (230) via the kernel (330).
[0110] In one embodiment, the DSP (230) may transmit a signal to the CPU (220) indicating the point in time of the PCM frames being played through the speaker while controlling the speaker (e.g., the speaker (240) of FIG. 2) using PCM frames. The DSP (230) may transmit the signal based on a request transmitted from the CPU (220). The embodiment is not limited thereto, and the DSP (230) may periodically (or repeatedly) transmit the signal to the CPU (220). The DSP (230) may inform the CPU (220) of the playback time of the audio frames decoded (or rendered) by the DSP (230). The signal may be received by the kernel (330) executed by the CPU (220). The player framework (320) executed by the CPU (220) may identify the point in time from the signal received via the kernel (330).
[0111] In one embodiment, the CPU (220) may use information returned from the DSP (230) (e.g., the playback times of audio frames decoded (or rendered) by the DSP (230)) to determine whether to perform synchronization of video frames and audio frames (e.g., operation (460) of FIG. 4). Hereinafter, with reference to FIG. 7, an exemplary operation of the CPU (220) that has determined to perform synchronization of video frames and audio frames to control playback of the video frames is described.
[0112] FIG. 7 illustrates the operation of an electronic device that adjusts the playback time of a video based at least on the duration of sound output from a speaker (e.g., speaker (240) of FIG. 2). The electronic device of FIG. 7 may include the electronic device (101) of FIG. 1 and / or FIG. 2. The operations of FIG. 7 may be performed by the AP (210) and / or CPU (220) of FIG. 2.
[0113] Referring to FIG. 7, video frames (e.g., the first video frame (710-1) to the fifth video frame (710-5) of FIG. 1) obtained by an electronic device performing demultiplexing on multimedia content are illustrated. The electronic device can perform decoding on the video frames to obtain images compressed in each of the video frames. The images may be referred to as image frames of the video. The electronic device can sequentially output the images through a display (e.g., the display (110) of FIG. 2) to provide a visual effect similar to playing the video.
[0114] Referring to FIG. 7, timestamps indicating the time at which images corresponding to each of the video frames are displayed are recorded within blocks corresponding to each of the video frames. For example, in a first video frame (710-1), an image to be displayed at 2100 ms may be compressed. For example, the electronic device may perform decoding on a second video frame (710-2) to obtain an image to be displayed at 2133 ms. For example, the electronic device may display the compressed image in the third video frame (710-3) on the display at a time point of 2166 ms after playing the video frames. For example, the fourth video frame (710-4) may include information (e.g., a header and / or metadata) indicating that the compressed image in the fourth video frame (710-4) is to be displayed at 2200 ms.
[0115] While sequentially outputting images corresponding to the video frames of FIG. 7, the electronic device may output audio frames multiplexed to the multimedia content. The output of the audio frames may be performed by another processor (e.g., the DSP (230) of FIG. 2) different from the processor that processes the video frames (e.g., the CPU (220) of FIG. 2). The other processor may be referred to as an offload device, as described above. When decoding and / or rendering audio frames using the DSP, the CPU connected to the DSP may obtain, from the DSP, the time (or timestamp) of the audio frame being decoded (or rendered) by the DSP.
[0116] For example, based on operation (440) of FIG. 4, it is assumed that the time difference between the expected playback time and the playback time of the audio frame exceeds a specified threshold value of 2200 ms. Within this assumption, the electronic device can perform operation (445) of FIG. 4 and store a timestamp indicating 2200 ms as the timestamp of operation (445). Within this assumption, the electronic device can identify whether the timestamp received from the DSP (e.g., the time of operation (455) of FIG. 4) exceeds 2200 ms. For example, the electronic device can execute the playback time management unit (257) to determine or determine whether the timestamp received from the DSP exceeds 2200 ms. If the timestamp received from the DSP indicates a time exceeding 2200 ms, the electronic device may change the time of the media clock and / or the media clock, which indicates the point in time of the video frame displayed on the display, to at least partially compensate for the time difference.
[0117] For example, if the electronic device adjusts the media clock and / or the time of the media clock by 30 ms, based on the change in the time of the media clock, the electronic device may display on the display a third video frame (710-3) (e.g., a video frame displayed in the time interval between 2166 ms and 2200 ms) corresponding to 2170 ms (= 2200 ms - 30 ms). Since the time of the media clock is decreased, the fourth video frame (710-4) after 2170 ms may be displayed relatively slowly. For example, since the time of the media clock is changed from 2200 ms (or a time exceeding 2200 ms) to 2170 ms, the electronic device may display a video frame corresponding to 2170 ms (e.g., the third video frame (710-3)) instead of a video frame corresponding to 2200 ms. In the above example, since the video frame corresponding to 2170 ms is displayed, there may be a delay in the time it takes for the video frames to be rendered.
[0118] In one embodiment, the electronic device may stop DSP-based decoding and / or rendering of audio frames when a parameter related to playback of video frames (e.g., a media clock and / or a time of the media clock) needs to be adjusted relatively significantly. For example, the electronic device may directly perform decoding and / or rendering of audio frames using a CPU. For example, the electronic device may directly perform decoding of audio frames using a CPU when an accumulated time difference (e.g., a time difference between an expected playback time and an actual playback time) of operation (440) of FIG. 4 exceeds a designated threshold value set for CPU-based decoding of audio frames (e.g., another designated threshold value that exceeds the threshold value of operation (440)). For example, the electronic device may control the DSP to stop DSP-based decoding and / or rendering of audio frames.
[0119] As described above with reference to FIG. 4, even if the time difference becomes excessively large, since the electronic device controls the playback of video frames using a limited value (e.g., a time adjustment value), the electronic device can stop the execution of the offload audio and perform CPU-based decoding and / or rendering of the audio frames among the CPU or DSP, thereby more quickly compensating for the mismatch caused by the execution of the offload audio.
[0120] The embodiment is not limited thereto, and if the calculation of the expected playback time based on the profile of the audio codec is not possible (e.g., if the audio frames are encoded in a size different from the size of the standardized audio codec), the electronic device may stop executing the offloaded audio and directly perform CPU-based decoding and / or rendering of the audio frames.
[0121] In one embodiment, performing off-load audio may be customized. For example, the electronic device may enable or disable off-load audio based on user input. When off-load audio is enabled and mismatches between video frames and audio frames occur frequently (or relatively significantly), the electronic device may display a UI (e.g., a pop-up window confirming the disabling of off-load audio) to receive an input to disable off-load audio. Upon receiving such an input via the UI, the electronic device may stop performing off-load audio and directly perform decoding and / or rendering of audio frames using the CPU.
[0122] As described above, according to one embodiment, the electronic device can compensate for or reduce mismatch of video frames and audio frames by using information transmitted from the DSP (e.g., a timestamp indicating the total time of audio frames decoded by the DSP) and a time difference between an expected playback time and an actual playback time predicted in a demultiplexing process of the audio frames. The electronic device can perform an operation to compensate for the mismatch when executing a software application that plays multimedia content that is a combination of video and audio, such as a video player, a gallery, and / or a streaming application.
[0123] FIG. 8 is a block diagram of an electronic device (801) within a network environment (800) according to various embodiments. Referring to FIG. 8, in the network environment (800), the electronic device (801) may communicate with the electronic device (802) via a first network (898) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (804) or the server (808) via a second network (899) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (801) may communicate with the electronic device (804) via the server (808). According to one embodiment, the electronic device (801) may include a processor (820), a memory (830), an input module (850), an audio output module (855), a display module (860), an audio module (870), a sensor module (876), an interface (877), a connection terminal (878), a haptic module (879), a camera module (880), a power management module (888), a battery (889), a communication module (890), a subscriber identification module (896), or an antenna module (897). In some embodiments, the electronic device (801) may omit at least one of these components (e.g., the connection terminal (878)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (876), the camera module (880), or the antenna module (897)) may be integrated into one component (e.g., the display module (860)).
[0124] The processor (820) may, for example, execute software (e.g., a program (840)) to control at least one other component (e.g., a hardware or software component) of the electronic device (801) connected to the processor (820) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (820) may store commands or data received from other components (e.g., a sensor module (876) or a communication module (890)) in a volatile memory (832), process the commands or data stored in the volatile memory (832), and store result data in a non-volatile memory (834). According to one embodiment, the processor (820) may include a main processor (821) (e.g., a central processing unit or an application processor) or an auxiliary processor (823) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (821). For example, when the electronic device (801) includes the main processor (821) and the auxiliary processor (823), the auxiliary processor (823) may be configured to use less power than the main processor (821) or to be specialized for a given function. The auxiliary processor (823) may be implemented separately from the main processor (821) or as a part thereof.
[0125] The auxiliary processor (823) may control at least a portion of functions or states associated with at least one component (e.g., a display module (860), a sensor module (876), or a communication module (890)) of the electronic device (801), for example, on behalf of the main processor (821) while the main processor (821) is in an inactive (e.g., sleep) state, or together with the main processor (821) while the main processor (821) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (823) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (880) or a communication module (890)). In one embodiment, the auxiliary processor (823) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (801) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (808)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0126] The memory (830) can store various data used by at least one component (e.g., the processor (820) or the sensor module (876)) of the electronic device (801). The data can include, for example, software (e.g., the program (840)) and input data or output data for commands related thereto. The memory (830) can include volatile memory (832) or non-volatile memory (834).
[0127] The program (840) may be stored as software in the memory (830) and may include, for example, an operating system (842), middleware (844), or an application (846).
[0128] The input module (850) can receive commands or data to be used in a component of the electronic device (801) (e.g., a processor (820)) from an external source (e.g., a user) of the electronic device (801). The input module (850) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0129] The audio output module (855) can output audio signals to the outside of the electronic device (801). The audio output module (855) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0130] The display module (860) can visually provide information to an external party (e.g., a user) of the electronic device (801). The display module (860) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (860) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0131] The audio module (870) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (870) can acquire sound through the input module (850), output sound through the sound output module (855), or an external electronic device (e.g., electronic device (802)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (801).
[0132] The sensor module (876) can detect the operating status (e.g., power or temperature) of the electronic device (801) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (876) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0133] The interface (877) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (801) with an external electronic device (e.g., the electronic device (802)). In one embodiment, the interface (877) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0134] The connection terminal (878) may include a connector through which the electronic device (801) may be physically connected to an external electronic device (e.g., the electronic device (802)). In one embodiment, the connection terminal (878) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0135] The haptic module (879) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (879) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0136] The camera module (880) can capture still images and videos. According to one embodiment, the camera module (880) may include one or more lenses, image sensors, image signal processors, or flashes.
[0137] The power management module (888) can manage the power supplied to the electronic device (801). According to one embodiment, the power management module (888) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0138] A battery (889) may power at least one component of the electronic device (801). In one embodiment, the battery (889) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0139] The communication module (890) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (801) and an external electronic device (e.g., electronic device (802), electronic device (804), or server (808)), and the performance of communication through the established communication channel. The communication module (890) may operate independently from the processor (820) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (890) may include a wireless communication module (892) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (894) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (804) via a first network (898) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (899) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (892) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (896) to verify or authenticate the electronic device (801) within a communication network such as the first network (898) or the second network (899).
[0140] The wireless communication module (892) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (892) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (892) may support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (892) may support various requirements specified in the electronic device (801), an external electronic device (e.g., the electronic device (804)), or a network system (e.g., the second network (899)). According to one embodiment, the wireless communication module (892) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0141] The antenna module (897) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (897) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (897) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (898) or the second network (899), may be selected from the plurality of antennas by, for example, the communication module (890). A signal or power may be transmitted or received between the communication module (890) and an external electronic device via the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (897).
[0142] According to various embodiments, the antenna module (897) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0143] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0144] According to one embodiment, commands or data may be transmitted or received between the electronic device (801) and an external electronic device (804) via a server (808) connected to a second network (899). Each of the external electronic devices (802 or 804) may be the same or a different type of device as the electronic device (801). According to one embodiment, all or part of the operations executed in the electronic device (801) may be executed in one or more of the external electronic devices (802, 804, or 808). For example, when the electronic device (801) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (801) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (801). The electronic device (801) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (801) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (804) may include an Internet of Things (IoT) device. The server (808) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (804) or the server (808) may be included in the second network (899).The electronic device (801) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0145] Fig. 9 is a block diagram (900) of an audio module (870) according to various embodiments. Referring to Fig. 9, the audio module (870) may include, for example, an audio input interface (910), an audio input mixer (920), an analog to digital converter (ADC) (930), an audio signal processor (940), a digital to analog converter (DAC) (950), an audio output mixer (960), or an audio output interface (970).
[0146] The audio input interface (910) can receive an audio signal corresponding to a sound acquired from outside the electronic device (801) as part of the input module (850) or through a microphone (e.g., a dynamic microphone, a condenser microphone, or a piezo microphone) configured separately from the electronic device (801). For example, when the audio signal is acquired from an external electronic device (802) (e.g., a headset or a microphone), the audio input interface (910) can receive the audio signal by being directly connected to the external electronic device (802) through a connection terminal (878) or wirelessly (e.g., Bluetooth communication) through a wireless communication module (892). According to one embodiment, the audio input interface (910) can receive a control signal (e.g., a volume control signal received through an input button) related to the audio signal acquired from the external electronic device (802). The audio input interface (910) includes a plurality of audio input channels and can receive different audio signals for each corresponding audio input channel among the plurality of audio input channels. According to one embodiment, additionally or alternatively, the audio input interface (910) can receive audio signals from other components of the electronic device (801), such as the processor (820) or the memory (830).
[0147] The audio input mixer (920) can synthesize a plurality of input audio signals into at least one audio signal. For example, according to one embodiment, the audio input mixer (920) can synthesize a plurality of analog audio signals input through the audio input interface (910) into at least one analog audio signal.
[0148] The ADC (930) can convert an analog audio signal into a digital audio signal. For example, according to one embodiment, the ADC (930) can convert an analog audio signal received through the audio input interface (910) or, additionally or alternatively, an analog audio signal synthesized through the audio input mixer (920) into a digital audio signal.
[0149] The audio signal processor (940) may perform various processing on a digital audio signal input through the ADC (930) or a digital audio signal received from another component of the electronic device (801). For example, according to one embodiment, the audio signal processor (940) may change a sampling rate, apply one or more filters, perform interpolation processing, amplify or attenuate all or part of a frequency band, process noise (e.g., noise or echo reduction), change a channel (e.g., switching between mono and stereo), mix, or extract a specified signal on one or more digital audio signals. According to one embodiment, one or more functions of the audio signal processor (940) may be implemented in the form of an equalizer.
[0150] The DAC (950) can convert a digital audio signal into an analog audio signal. For example, according to one embodiment, the DAC (950) can convert a digital audio signal processed by an audio signal processor (940) or a digital audio signal obtained from another component of the electronic device (801) (e.g., a processor (820) or a memory (830)) into an analog audio signal.
[0151] The audio output mixer (960) can synthesize a plurality of audio signals to be output into at least one audio signal. For example, according to one embodiment, the audio output mixer (960) can synthesize an audio signal converted into analog through the DAC (950) and another analog audio signal (e.g., an analog audio signal received through the audio input interface (910)) into at least one analog audio signal.
[0152] The audio output interface (970) can output an analog audio signal converted by the DAC (950), or additionally or alternatively, an analog audio signal synthesized by the audio output mixer (960), to the outside of the electronic device (801) through the audio output module (855). The audio output module (855) can include, for example, a speaker, such as a dynamic driver or a balanced armature driver, or a receiver. According to one embodiment, the audio output module (855) can include a plurality of speakers. In this case, the audio output interface (970) can output an audio signal having a plurality of different channels (e.g., stereo or 5.1 channels) through at least some of the speakers. According to one embodiment, the audio output interface (970) can be directly connected to an external electronic device (802) (e.g., an external speaker or a headset) through a connection terminal (878) or wirelessly through a wireless communication module (892) to output an audio signal.
[0153] According to one embodiment, the audio module (870) can generate at least one digital audio signal by synthesizing a plurality of digital audio signals using at least one function of the audio signal processor (940) without separately having an audio input mixer (920) or an audio output mixer (960).
[0154] According to one embodiment, the audio module (870) may include an audio amplifier (not shown) (e.g., a speaker amplifier circuit) capable of amplifying an analog audio signal input through the audio input interface (910) or an audio signal to be output through the audio output interface (970). According to one embodiment, the audio amplifier may be configured as a separate module from the audio module (870).
[0155] The electronic device (801) of FIGS. 8 and 9 may be an example of the electronic device (101) of FIGS. 1 and 2. The DSP (230) of FIG. 2 may include at least a portion of the audio module (870) described above with reference to FIGS. 8 and 9.
[0156] In one embodiment, a method may be required to synchronously output video and audio of multimedia content using different processors (or circuits). In one embodiment, a method may be required to identify and / or predict mismatches between video and audio of multimedia content decoded by different processors. In one embodiment, a method may be required to control playback of video and / or audio based on the identified mismatches. As described above, according to one embodiment, an electronic device may include a display, a speaker, one or more storage media, a memory storing instructions, a first processor including processing circuitry, and a second processor including processing circuitry. The instructions, when executed by the first processor, may cause the electronic device to receive an input for playing back multimedia content. The instructions, when executed by the first processor, may cause the electronic device, in response to the input, to identify an audio codec associated with audio frames included in the multimedia content. The instructions, when executed by the first processor, may cause the electronic device to obtain a first set of audio frames from the multimedia content. The instructions, when executed by the first processor, may cause the electronic device to identify playback times of the audio frames included in the first state. The instructions, when executed by the first processor, may cause the electronic device to combine time differences of each of the playback times with respect to an expected playback time of one audio frame indicated by the audio codec, and to obtain a first point in time having a time difference exceeding a specified reference value.The instructions, when executed by the first processor, may cause the electronic device to transmit the first set of audio frames to the second processor. The instructions, when executed by the second processor, may cause the electronic device to decode the audio frames based on receiving the first set of audio frames. The instructions, when executed by the second processor, may cause the electronic device to transmit, to the first processor, a second point in time of the PCM frames to be played back through the speaker based on controlling the speaker using PCM frames obtained based on the decoding, each corresponding to the audio frames. The instructions, when executed by the first processor, may cause the electronic device to control playback of the video through the display to at least partially compensate for the time difference based on receiving a second point in time after the first point in time from the second processor.
[0157] For example, the instructions, when executed by the first processor, may cause the electronic device to determine the playback times of the audio frames based on the PCM sample sizes and sampling rates of each of the audio frames.
[0158] For example, the instructions, when executed by the first processor, may cause the electronic device to determine the expected playback time based on a PCM sample size per audio frame defined for the audio codec.
[0159] For example, the instructions, when executed by the first processor, may cause the electronic device to transmit a signal to the second processor requesting transmission of the second point in time while controlling the display to play the video.
[0160] For example, the instructions, when executed by the second processor, may cause the electronic device to periodically transmit signals to the first processor indicating a point in time of the PCM frames being played through the speaker, including the second point in time, while controlling the speaker using the PCM frames.
[0161] For example, the instructions, when executed by the first processor, may cause the electronic device to change the time of a media clock, which represents a point in time of a video frame displayed on the display, to at least partially compensate for the time difference.
[0162] For example, the instructions, when executed by the first processor, may cause the electronic device to control the second processor to stop performing decoding of the audio frames based on the second processor, based on obtaining the first point in time having a time difference that exceeds the specified reference value, which exceeds another specified reference value. The instructions, when executed by the first processor, may cause the electronic device to directly perform decoding of the audio frames.
[0163] For example, the instructions, when executed by the first processor, may cause the electronic device to perform demultiplexing on the multimedia content to obtain the first set of audio frames from the multimedia content.
[0164] For example, the instructions, when executed by the first processor, may cause the electronic device to obtain, from the multimedia content, the first set of audio frames, the first set having a size of a buffer allocated within the memory for the demultiplexing.
[0165] For example, the electronic device may include an application processor (AP). The AP may include the first processor, which includes a core circuit of a central processing unit (CPU), and the second processor, which is a digital signal processor (DSP) configured to perform digital signal processing on the audio frames.
[0166] As described above, in one embodiment, a method of an electronic device may be provided. The electronic device may include a display, a speaker, a first processor, and a second processor. The method may include an operation of controlling the first processor to receive an input for playing multimedia content. The method may include an operation of controlling the first processor to, in response to the input, identify an audio codec associated with audio frames included in the multimedia content. The method may include an operation of controlling the first processor to obtain a first set of audio frames from the multimedia content. The method may include an operation of controlling the first processor to identify playback times of the audio frames included in the first set. The method may include an operation of controlling the first processor to combine time differences of each of the playback times with respect to an expected playback time of one audio frame indicated by the audio codec, and to obtain a first point in time having a time difference exceeding a specified reference value. The method may include controlling the first processor to transmit the first set of audio frames to the second processor. The method may include controlling the display to play a video within the multimedia content included in a time interval corresponding to the first set, by controlling the first processor. The method may include controlling the second processor to perform decoding on the audio frames based on receiving the first set of audio frames.The method may include an operation of controlling the second processor to control the speaker using pulse code modulation (PCM) frames, each of which is obtained based on the decoding and corresponding to the audio frames, and transmitting a second point in time of the PCM frames to be played back through the speaker to the first processor. The method may include an operation of controlling the first processor to receive a second point in time after the first point in time from the second processor, and controlling the playback of the video through the display to at least partially compensate for the time difference.
[0167] For example, the act of identifying the playback times may include the act of determining the playback times of the audio frames based on the PCM sample sizes and sampling rates of each of the audio frames.
[0168] For example, the act of identifying the audio codec may include the act of determining the expected playback time based on a PCM sample size per audio frame defined by the audio codec.
[0169] For example, the operation of controlling the display may include an operation of transmitting a signal requesting transmission of the second point in time to the second processor while controlling the display to play the video.
[0170] For example, the operation of transmitting the second point in time may include an operation of periodically transmitting, to the first processor, signals representing the point in time of the PCM frames played through the speaker, including the second point in time, while controlling the speaker using the PCM frames.
[0171] For example, the act of controlling playback of the video may include changing the time of a media clock, which indicates the point in time of a video frame displayed on the display, to at least partially compensate for the time difference.
[0172] For example, the operation of acquiring the first point in time may include controlling the second processor to stop performing decoding on the audio frames based on acquiring the first point in time having a time difference exceeding the specified reference value or exceeding another specified reference value. The operation of acquiring the first point in time may include directly performing decoding on the audio frames.
[0173] For example, the operation of obtaining the first set may include an operation of obtaining the first set of audio frames from the multimedia content by performing demultiplexing on the multimedia content.
[0174] For example, the operation of acquiring the first set may include an operation of acquiring the audio frames of the first set, having a size of a buffer allocated in the memory for the demultiplexing, from the multimedia content.
[0175] As described above, in one embodiment, a non-transitory computer-readable storage medium including instructions may be provided. The instructions, when executed by an electronic device including a display, a speaker, a first processor, and a second processor, may cause the electronic device to control the first processor to detect an input for playing multimedia content. The instructions, when executed by the electronic device, may cause the electronic device to check an audio codec of the multimedia content in response to the input. The instructions, when executed by the electronic device, may cause the electronic device to identify a difference between an expected playback time of audio frames included in the multimedia content, indicated by the checked audio codec, and a time at which an audio signal indicated by the audio frames is output from the speaker based on decoding of the audio frames based on the second processor. The instructions, when executed by the electronic device, may cause the electronic device to change the time of the video displayed on the display based on identifying the difference while controlling the display to play the video of the multimedia content.
[0176] According to one embodiment of the present invention, an electronic device as described above may include a first processor including a display, a speaker, and processing circuitry, and a second processor including processing circuitry. The second processor may be configured to decode a set of audio frames based on receiving the set of audio frames from the first processor. The second processor may be configured to control a speaker such that an acoustic signal represented by the sequence of PCM frames is output through the speaker based on identifying a sequence of PCM frames corresponding to each of the audio frames based on the decoding. The second processor may be configured to transmit, to the first processor, information indicating one PCM frame, among the PCM frames, corresponding to the acoustic signal being reproduced through the speaker while controlling the speaker based at least on the sequence of PCM frames. The above information may be used to control playback of video frames being output through the display controlled by the first processor based on a difference between an expected playback time obtained by decoding the audio frames and a playback time indicated by the information.
[0177] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0178] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0179] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0180] Various embodiments of the present document may be implemented as software (e.g., a program (840)) including one or more instructions stored in a storage medium (e.g., an internal memory (836) or an external memory (838)) readable by a machine (e.g., an electronic device (801)). For example, a processor (e.g., a processor (820)) of the machine (e.g., an electronic device (801)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0181] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0182] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0183] As used herein, the term "if" will be understood to mean "when, upon," "in response to determining," or "in response to detecting," depending on the context. Similarly, "if it is determined to," or "if [the stated condition or event] is detected," will optionally be understood to mean "upon determining," or "in response to determining," "upon detecting [the stated condition or event]," or "in response to detecting [the stated condition or event]."
[0184] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0185] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.
[0186] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.
[0187] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above description. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0188] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.
Claims
In electronic devices, display; speaker; A memory comprising one or more storage media for storing instructions; a first processor comprising a processing circuit; and a second processor including a processing circuit, The above instructions, when executed by the first processor, cause the electronic device to: Receive input for playing multimedia content; In response to the above input, identifying an audio codec associated with audio frames included in the multimedia content; Obtaining a first set of audio frames from the multimedia content; Identifying the playback times of the audio frames included in the first set; By combining the time differences of each of the above playback times with respect to the expected playback time of one audio frame indicated by the above audio codec, a first point in time having a time difference exceeding a specified reference value is obtained; and transmitting said first set of said audio frames to said second processor; Causing the display to be controlled to play a video within the multimedia content, which is included in a time interval corresponding to the first set, and The above instructions, when executed by the second processor, cause the electronic device to: performing decoding on said audio frames based on receiving said first set of said audio frames; and Based on the decoding, the speaker is controlled using PCM (pulse code modulation) frames, each corresponding to the audio frames, and the first processor causes the second point in time of the PCM frames to be played back through the speaker, and The above instructions, when executed by the first processor, cause the electronic device to: Controlling playback of the video through the display to at least partially compensate for the time difference based on receiving a second point in time after the first point in time from the second processor, Electronic devices. In claim 1, The above instructions, when executed by the first processor, cause the electronic device to: Causing the playback times of the audio frames to be determined based on the PCM sample sizes and sampling rates of each of the audio frames. Electronic devices. In claim 1, The above instructions, when executed by the first processor, cause the electronic device to: Causing the expected playback time to be determined based on the PCM sample size per audio frame defined for the above audio codec. Electronic devices. In claim 1, The above instructions, when executed by the first processor, cause the electronic device to: While controlling the display to play the video, causing the second processor to transmit a signal requesting transmission of the second point in time. Electronic devices. In claim 1, The above instructions, when executed by the second processor, cause the electronic device to: While controlling the speaker using the PCM frames, causing the first processor to periodically transmit signals indicating the time points of the PCM frames played through the speaker, including the second time point. Electronic devices. In claim 1, The above instructions, when executed by the first processor, cause the electronic device to: To at least partially compensate for the above time difference, causing the time of the media clock, which indicates the point in time of the video frame displayed on the display, to be changed, Electronic devices. In claim 1, The above instructions, when executed by the first processor, cause the electronic device to: Controlling the second processor to stop performing decoding on the audio frames based on the second processor based on obtaining the first point in time having a time difference exceeding the specified reference value and exceeding another specified reference value; causing the decoding of the above audio frames to be performed directly, Electronic devices. In claim 1, The above instructions, when executed by the first processor, cause the electronic device to: performing demultiplexing on the multimedia content, thereby causing the first set of audio frames to be obtained from the multimedia content; Electronic devices. In claim 8, The above instructions, when executed by the first processor, cause the electronic device to: Causing the first set of audio frames, having a size of a buffer allocated within the memory for the demultiplexing, to be acquired from the multimedia content. Electronic devices. In claim 1, Further comprising an AP (application processor), wherein the AP: The first processor including a core circuit of a CPU (central processing unit); and The second processor, which is a digital signal processor (DSP) configured to perform digital signal processing on the audio frames including, Electronic devices. In a method of an electronic device, the electronic device includes a display, a speaker, a first processor, and a second processor. Controlling the above first processor: An action to receive input for playing multimedia content; In response to the above input, an operation of identifying an audio codec associated with audio frames included in the multimedia content; An operation of obtaining a first set of audio frames from the multimedia content; An operation of identifying the playback times of the audio frames included in the first set; An operation of combining the time differences of each of the above playback times with respect to the expected playback time of one audio frame indicated by the above audio codec, and obtaining a first point in time having a time difference exceeding a specified reference value; an operation of transmitting said first set of said audio frames to said second processor; and An operation for controlling the display to play a video within the multimedia content included in the time interval corresponding to the first set, and Controlling the second processor: An operation of performing decoding on said audio frames based on receiving said first set of said audio frames; and An operation of transmitting a second point in time of the PCM frames played through the speaker to the first processor based on controlling the speaker using PCM (pulse code modulation) frames obtained based on the decoding and each corresponding to the audio frames, and Controlling the above first processor: An operation of controlling playback of the video through the display to at least partially compensate for the time difference based on receiving a second point in time after the first point in time from the second processor, method. In claim 11, the operation of identifying the playback times comprises: An operation of determining the playback times of the audio frames based on the PCM sample sizes and sampling rates of each of the audio frames, method. In claim 11, the operation of identifying the audio codec comprises: An operation for determining the expected playback time based on the PCM sample size per audio frame defined by the audio codec, method. In claim 11, the operation of controlling the display comprises: An operation of transmitting a signal requesting transmission of the second point in time to the second processor while controlling the display to play the video, method. In claim 11, the operation of transmitting the second point in time comprises: While controlling the speaker using the PCM frames, the first processor includes an operation of periodically transmitting signals indicating the time points of the PCM frames played through the speaker, including the second time point. method.
Citation Information
Patent Citations
Wet dust removal or gas condensation system using multi-layer porous mesh panels
KR1020240039974A
Men's Protrusion Prevention Device
KR1020240042590A
Plating solution of negative electrode for lithium secondary battery, negative electrode including the same, manufacturing method of lithium secondary battery
KR1020240168014A
Heat transfer apparatus and rechargeable battery comprising the same
KR1020250118108A
Laser welding torch with lens protection
KR102883411B1