Methods and systems for time alignment in multi-channel encoding
By adapting coding frame lengths and using resampling to align sampling frequencies, the method addresses time alignment issues in multi-channel encoding, enabling efficient joint encoding and decoding of signals with varying frequencies.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- DOLBY INTERNATIONAL AB
- Filing Date
- 2025-12-01
- Publication Date
- 2026-06-11
AI Technical Summary
Multi-channel signals with different sampling frequencies face challenges in maintaining time alignment during encoding, particularly in applications like biomedical signal analysis, where joint encoding is desired but constrained by processing frame lengths that do not match the varying sampling frequencies.
Adapt the coding frame length per channel based on the sampling frequency to maintain time alignment, using resampling processes to adjust sampling frequencies to compatible transform lengths, and introduce frame multipliers to ensure synchronization across channels.
Ensures time alignment of multi-channel signals by selecting appropriate processing frame lengths and resampling, facilitating efficient joint encoding and decoding, especially in biomedical applications.
Smart Images

Figure EP2025085010_11062026_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR TIME ALIGNMENT IN MULTI-CHANNEL ENCODINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority from United States Provisional Patent Application No. 63 / 726,812, filed on 2 December 2024, and from European patent application EP 25 183 840.5, filed on 18 June 2025, both of which are incorporated by reference herein in their entirety.TECHNICAL FIELD OF THE INVENTION
[0002] The present invention relates to a multi-channel encoding, and more specifically situations where different input channels have been sampled with different sampling frequencies.BACKGROUND OF THE INVENTION
[0003] In some signal coding applications, the sampling frequency (Fs) of the signals can vary widely, e.g., from single digit Hertz up to tens of kHz. This creates challenges, especially when information from several different sources is acquired with different sampling frequencies and included in one signal including one channel for each source. Such a signal is herein referred to as a multi-channel signal.
[0004] The situation where a multichannel input signal may contain individual channels sampled at different sampling frequencies can be encountered in many scenarios for acquisition of biomedical signals. For example, during the acquisition one may use multiple sensing devices, each of them providing a subset of channels (e.g., EEG channels). The sensing devices may be configured or limited to specific sampling frequencies. The channels may be time aligned at the input of the encoder, but their sampling frequencies may differ. It may be still desired to be able to encode these channels jointly, for example, when the channels are supposed to be analyzed jointly after the decoding.
[0005] In other applications it may be beneficial to encode jointly different types of medical signals. For example, one can envision a scenario, where a single channel PPG signal is combined with a single channel or multiple channel ECG signal. Again, the channels may be time aligned at the input of the encoder, but their sampling frequencies may differ. It may be still beneficial to allow for joint coding, if on the decoder side, an algorithm is going to analyze the decoded signal to estimate a parameter of interest (e.g., in the case of PPG and ECG acquisition such a joint analysis on the decoder side could for example be an estimation of pulse wave velocity, cuffless measurement of blood pressure, etc.)
[0006] Joint encoding of multi-channel signals where different channels have different sampling frequencies may occur also in other areas.GENERAL DISCLOSURE OF THE INVENTION
[0007] For efficiency reasons, an encoder process typically introduces one or several constraints on a processing frame length used in the encoder (and typically also in the decoder). For example, when a frequency transform is applied in the encoder, the processing frame should correspond to a restricted set of appropriate transform lengths (e.g. 512, 768, 1024). In a situation where a multi-channel input is encoded, this may cause different encoded channels to diverge over time unless corrected in some manner. Such time divergence is undesirable, especially in cases where analysis of the information in the channels require time alignment. This is clearly the case in the biomedical field, where a video feed of a patient or an organ needs to be in sync with e.g., an EEG and / or ECG of the patient.
[0008] The present disclosure seeks to mitigate this problem by adapting the coding frame length per channel, to compensate for different sampling frequencies in different channels.
[0009] According to a first aspect there is provided a method for encoding a multi-channel input signal, comprising receiving at least two input channels, xn, each being acquired with a different sampling frequency, fsn, for each input channel, selecting a processing (coding) frame length, Ln, and encoding each input channel using the processing frame length, wherein the encoding of each input signal involves a time-to-frequency transform, wherein the processing frame length for each channel is selected from a set of available transform lengths which are compatible with the respective time-to-frequency transform, based on the sampling frequency of each respective input channel, in order to maintain time alignment of the input channels.
[0010] With this approach, a different (allowable) processing frame length may be used in the encoding of different channels. The selection of an appropriate processing frame length (from a set of allowable lengths) is made based on the sampling frequency of the particular channel (or group of channels), in order to maintain time alignment of the different channels.
[0011] Depending on the various sampling frequencies and set of allowable processing frame lengths, it is not always possible to match each input channel with an appropriate transform length while maintaining time alignment. For this purpose, a resampling process may be implemented to ensure an appropriate relationship (e.g., ratio) between (resampled) sampling frequencies of different channels. The resampling may be achieved in various ways, for example by buffering the input channel in a buffer of appropriate length and sampling the buffer with a desired sampling frequency.
[0012] In some implementations, the resampling can be done so as to ensure a ratio between sampling frequencies (after resampling) that corresponds to a ratio between two allowable processing frame lengths. By selecting these processing frame lengths, time alignment may be maintained.
[0013] In other implementations, it may be appropriate to introduce a frame multiplier for each channel, and forming a channel package including a number of processing frames indicated by the frame multiplier. The resampling can then be done so as to ensure a ratio between sampling frequencies (after resampling) that corresponds to a ratio between two allowable processing frame lengths multiplied by the respective frame multiplier. By selecting these processing frame lengths, time alignment may be maintained.
[0014] The constraints on processing frame length may be caused by different aspects of the encoding process.
[0015] In other implementations, the encoding is configured to operate on time domain frames having one - or one of several - predefined time domain frame lengths. In this case, the processing frame length should be selected to correspond to one of these predetermined time domain frame lengths.
[0016] In a specific implementation, the encoding involves frequency domain prediction residual encoding. An example of such encoding is disclosed in US patent application No. 63 / 713,216 and European patent application 24209441.5, both herewith incorporated by reference. In this case, a prediction order applied when encoding an input channel may be selected based on the sampling frequency of the input channel.
[0017] On the decoder side, it may be desirable to resample at least one of the decoded channels based on side information relating to the resampling performed on the encoder side. Such side information could be included in the bitstream together with the encoded signal. Such resampling would then recreate the original input channels, and may be advantageous if the channels are processed by systems that are configured based on a specific sampling frequency. On the other hand, in other situations, it may be more advantageous to maintain the resampled version of the channels, as this facilitates continued time alignment.
[0018] A further aspect relates to a system for encoding a multi-channel biomedical input signal having at least two input channels, xn, each being acquired with a different sampling frequency, fsn, comprising a coding adaptation block configured to, for each input signal, select a processing frame length, Ln, and a set of encoders, each configured to encode one of the input channel using the processing frame length, wherein the encoders are configured to apply a time- to-frequency transform to each input signal, wherein the coding adaptation block is configured to select the processing frame length for each channel from a set of available transform lengthswhich are compatible with the respective time-to-frequency transform, based on the sampling frequency of each respective input channel, in order to maintain time alignment of the input channels.
[0019] Other aspects relate to: an encoder comprising a processor and a memory storing instructions configured to, when executed by the processor, cause the processor to perform the method of the first aspect, a computer program product comprising instructions configured to, when executed by a processor, cause the processor to perform the method according to the first aspect, and a non-transitory computer readable medium comprising instructions configured to, when executed by a processor, cause the processor to perform the method according to the first aspect.BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The present invention will be described in more detail with reference to the appended drawings, showing currently preferred embodiments of the invention.
[0021] Figure 1 shows a coding architecture according to an embodiment of the present invention.
[0022] Figure 2 shows an example encoding process to illustrate resampling and framing according to the embodiment in figure 1.
[0023] Figure 3 shows a coding architecture according to another embodiment of the present invention.
[0024] Figure 4 shows an example encoding process to illustrate resampling and framing according to the embodiment in figure 3.
[0025] Figure 5 shows a coding architecture according to yet another embodiment of the present invention.DETAILED DESCRIPTION OF CURRENTLY PREFERRED EMBODIMENTS
[0026] Embodiments of the present invention will be described in the following with reference to examples of encoding and decoding architecture. It is noted that the invention is not limited to this particular codec architecture. On the contrary, framing according to embodiments of the present invention may be useful for encoding and decoding in most codec systems.
[0027] Systems and methods disclosed in the present application may be implemented as software, firmware, hardware or a combination thereof. In a hardware implementation, the division of tasks does not necessarily correspond to the division into physical units; to thecontrary, one physical component may have multiple functionalities, and one task may be carried out by several physical components in cooperation.
[0028] The computer hardware may for example be a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular telephone, a smartphone, a web appliance, a network router, switch or bridge, or any machine capable of executing instructions (sequential or otherwise) that specify actions to be taken by that computer hardware. Further, the present disclosure shall relate to any collection of computer hardware that individually or jointly execute instructions to perform any one or more of the concepts discussed herein.
[0029] Certain or all components may be implemented by one or more processors that accept computer-readable (also called machine-readable) code containing a set of instructions that when executed by one or more of the processors carry out at least one of the methods described herein. Any processor capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken are included. Thus, one example is a typical processing system (i.e. a computer hardware) that includes one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system further may include a memory subsystem including a hard drive, SSD, RAM and / or ROM. A bus subsystem may be included for communicating between the components. The software may reside in the memory subsystem and / or within the processor during execution thereof by the computer system.
[0030] The one or more processors may operate as a standalone device or may be connected, e.g., networked to other processor(s). Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
[0031] The software may be distributed on computer readable media, which may comprise computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to a person skilled in the art, the term computer storage media includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, physical (non-transitory) storage media in various forms, such as EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Further, it is well known to the skilled person that communication media(transitory) typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.
[0032] In the following disclosure, at least one channel of the multi-channel input signal may be a biomedical signal. A biomedical signal in the context of the present disclosure may be a signal that relates to physiological information, and the signal may be electrical, physical or biochemical. The biomedical signal may relate to biological systems and conditions, examples of which may include Electrocardiography (ECG) data, Electroencephalography (EEG) data, Electromyography (EMG) data, and Photoplethysmogram (PPG) data, or signals for blood sugar level, heart rate, body temperature, respiratory rate and oxygen saturation. Further, a biomedical signal may comprise one or more channels of time domain biomedical signal samples. In other examples, the biomedical signals may relate to muscle and / or skin measurements. Any other medical signal and / or physical response would also be understood to be comprised by this definition.
[0033] Figure 1 shows an encoding architecture 10 receiving a multi-channel input signal, i.e. an input signal with several input channels. The fact that sampling frequencies of the input channels are different introduces a difficulty of maintaining the time-alignment between frames of the channels coming from subsets of channels operating at different sampling frequencies. Maintaining such an alignment is beneficial, since it allows for easy segmentation of the bitstream (e.g., a segment of decoded signal can be then easily selected from a bitstream, and the individual decoded channels are still aligned. This alignment is not trivial to achieve in a coding architecture which involves components that introduce an additional constraint on the frame size. An example of such component may be a frequency transform (e.g., DCT, MDCT, etc.), which may be computationally efficient only when the transform length is chosen appropriately. For example, the transform length may be a power of 2 (e.g., 512, 1024, 2048, etc. samples). In some situations, a transform length which is a multiple of a power of two is used, such as 768 (=3 x 256). This means that the framing of the input channel becomes constrained and it is no longer trivial to maintain frame alignment between subsets of channels sampled at different sampling frequencies. On the other hand, if sampling frequencies happened to correspond to the available transform lengths, maintaining the alignment would be easy.
[0034] For example, i the first set of channels may be sampled at fsi = 256 Hz, and the second set sampled at fs2 = 512 Hz. In the illustrated example, the objective is to generate frames of two seconds, and then apply transform coding to them using DCT of available transform length. In this case, a DCT of length of 512 could be used for the first set of channels and a DCT of length of 1024 for the second set of channels. Then the frames of the first set of channels arealways aligned with the frames of the second set of channels. Note, that while the available lengths of the transforms are typically constrained, we cannot expect that the input sampling frequencies will match these constraints in practical situations. In this case, we need a coding architecture that facilitates achieving signal alignment for the case where the input sampling frequencies are arbitrary.
[0035] The encoding architecture 10 in figure 1 includes two encoding paths 11, 12. A first set of input signals (channels), acquired with a first sampling frequency fsi are directed to the first encoding path, while a second set of input signals (channels), acquired with a second sampling frequency, fs2 are directed to the second encoding path. In the following, for simplicity, only one channel in each path is discussed.
[0036] Each encoding path here includes an encoder 13 followed by a quantization and entropy coding block 14. The encoder 13 is preceded by a framing block 15, configured to divide each input channel into a sequence of processing frames having frame lengths appropriate for the encoder 13. As mentioned, there may be various constraints on the choice of processing frame length.
[0037] For time domain processing, one or several predefined processing (encoder) frame lengths may be required by an encoding syntax. For example, if block switched coding is implemented, the algorithm for selecting block size sequence may require a predefined initial block size (sometimes referred to as “super” frame) which is then sub-divided appropriately. Examples of such algorithms may be found in co-pending US patent application titled “METHODS AND SYSTEMS FOR BLOCK SWITCHED CODING”, filed by Dolby International AB on December 2, 2024, naming Jonas Samuelsson as inventor and having Attorney Docket Number D24190USP1, which is herewith incorporated by reference in its entirety.
[0038] For frequency domain processing, one or several predefined transform lengths may be required by a time-frequency transform, such as a fast Fourier transform or variants thereof (DCT, DST, etc).
[0039] The architecture 10 further includes a coding adaptation block 16, configured to receive information about the sampling frequencies fsi and fs2. The information may e.g., be provided during system setup or be detected in the input channels.
[0040] Based on the sampling frequencies and constraints on the processing frame length, the coding adaptation block provides processing frame lengths Li L2 to the framing block 15 in each path.
[0041] In order to ensure time alignment between the channels, each processing frame should correspond to the same time period. Therefore, if a ratio between the samplingfrequencies in the first path 11 and the sampling frequency in the second path 12 is R, then the ratio between processing frame lengths should be 1 / R. In an ideal situation, there is a selection of processing frame lengths that perfectly matches the ratio R between sampling frequencies. For example, if the ratio R is two (e.g. fsi=500 Hz, fs2 = 1 kHz), and the encoder allows processing frames which are different by a factor of two (e.g. 512 and 1024 samples), then the path receiving the channel with higher sampling frequency (e.g. 1 kHz) is set to operate with the shorter processing frame length (e.g. 512 samples), while the path receiving the channel with lower sampling frequency (e.g. 500 Hz) is set to operate with the longer processing frame length (e.g. 1024 samples).
[0042] However, in a practical situation, the ratio R between sampling frequencies will not correspond to any available ratio between processing frame lengths. For this purpose, the paths 11 and 12 here include a resampling block 17, configured to resample the input channels. The code adaptation block 16 will determine a target ratio, i.e., a ratio between two available processing frame lengths, which is possible to obtain by resampling one or both of the input channels, and provide resampling factors to the resampling blocks 17 in one or both paths 11, 12. This process may be optimized based on various criteria. For example, it may be desirable to resample as few channels as possible. Or it may be desirable to minimize the largest change (increase or decrease) in sampling frequency. In other words, to facilitate coding efficiency, it may be desirable to keep the frequency of a resampled channels as close as possible to the sampling frequency of the original input channel. Further, a reduction in sampling frequency (down-sampling) may cause loss of information in the channel, and should therefore typically be avoided if possible, even if it means that additional channels need to be resampled (see example below).
[0043] The resampling blocks 17 may operate in various ways known in the art. For example, the input channel may be buffered in a buffer, and the buffer sampled with the desired sampling frequency. Depending on the original sampling rate, modified sampling rate, and the processing frame length, some boundary issues may arise between processing frames. However, this will not have any significant impact on the encoding process.
[0044] After resampling, a more suitable ratio R’ between the (resampled) sampling frequencies is achieved, allowing a selection of processing frame length in each path to ensure time alignment.
[0045] An example embodiment is depicted in figure 2. A 2-channel signal with sampling frequencies 3 Hz and 7 Hz in channel 1 and 2 respectively is encoded and bitstream packages are created such that the data in each packet allows decoding of both channels such that the decoded samples of each packet are time synchronized (time aligned) across the channels. In thisexample, the codec offers encoding using two block sizes, 64 or 128 samples, which notably differ by a factor of two. The resampling ratios are chosen such that sampling frequencies at the output of the resampler are Fsl ’ and Fs2’. Fsl ’ and Fs2’ are chosen such that the ratio between them is either one or equal to the ratio between the codec block sizes, i.e. two in this case, where a choice of one implies that channel 1 and channel 2 are encoded using the same codec block size, and a choice of two implies that channel 1 is encoded using the 64 samples block size (assuming Fsl’ < Fs2’) and that channel 2 is encoded using the 128 samples block size. In this example it is appropriate to set Fs2’ as two times Fsl’. A simple option, requiring only one resampling, could be to down-sample channel 2 to set Fs2’ = 6 Hz. However, as down-sampling comes at the cost of potentially losing information, in this case it is likely better to resample both channels, e.g. to set Fsl’ = 4 Hz and Fs2’ = 8 Hz. In other words, resampling ratios of 4:3 and 8:7 for channel 1 and channel 2, respectively.
[0046] On the decoder side (not depicted) the decoded channels can optionally be resampled with ratios 3:4 and 7:8 respectively to output a signal where the channels have the same sampling rates as the original input signal channels.
[0047] It is noted that with a limited number of allowable processing (encoder) frame lengths, and a large difference in sampling frequencies, the required resampling may become very significant. In such situations, it may be beneficial to also let the coding adaptation block 16 set a frame multiplier for each channel. Further, a channel package, i.e. a package output from each channel and collected in a bitstream package, will contain a number of processing frames indicated by the multiplier. The multiplier are set by the coding adaptation block such that each channel package represents an equal time duration.
[0048] Turning to figure 3, a more generalized embodiment is shown, where the input signal has N channels with sampling frequencies {fS1,fs2>andaset of K allowable encoder (codec) frame lengths {L1(L2, ... , LK}. The coding adaptation takes as inputs the set of sampling frequencies and the set of allowable encoder (processing) frame lengths, and determines a set of resampling ratios {Q1(Q2, ... , QN}> one ratio for each channel, a set of encoder frame lengths {Z1(l2, ... , lN], one encoder frame length for each channel and chosen from the set of allowable frame lengths, and a set of frame multipliers {u^ , u2, ... , uw], one frame multiplier for each channel. The frame multipliers are integers larger than zero and introduce freedom to group different number of encoder frames within each channel. The expression channel package is introduced as the samples in a channel that gets encoded and packaged together with channel packages from other channels to form a bitstream package. The length of a channel package is denoted Pn, and. Pn= unln, simply meaning that one or more encoder frames are encoded when constructing a package. At the output of the resampler of channel n the sampling frequency isFSn= Qnfsn- The number of seconds corresponding to a channel package is Pn / FSn. In order to maintain time alignment, the number of samples that are encoded in each channel package must correspond to the same number of seconds for all channels. Thus, to maintain time alignment, the coding adaptation determines resampling ratios Qn, frame length multipliers un, and encoder frame lengths lnsuch that the following is satisfied,or, more compactly,where c is a constant.
[0049] An example embodiment is depicted in figure 4 where fS1= 300 Hz, fS2= 1000 Hz, and fS3= 17 Hz. The allowable encoder frame lengths are= 512, L2= 768, and L3= 1024. There is considerable freedom in the coding adaptation. In this example, the coding adaptation determines resampling ratios Q1= 4 / 3, Q2= 1, and Q3= 20 / 17, encoder frame lengths 1- = 1024, l2= 768, and l3= 512, and frame multipliers u = 30, u2= 100, and u3= 3. This yields resampler output sampling frequencies FS1= 400 Hz, FS2= 1000 Hz, and Fs-3= 20 Hz. The number of samples in channel l’s package frame is P = 30x1024 = 30720, and for channel 2 and 3, P2= 100x768 = 76800 and P3= 3x512 = 1536. The number of seconds corresponding to channel l’s package frame is 30720 / 400 = 76.8s, and for channel 2 and 3, 76800 / 1000=76.8s and 1536 / 20 = 76.8s, respectively.
[0050] Figure 5 shows a multi-channel encoding architecture 20, where the encoder 23 is of the kind including a time-frequency transform 23a, in this case a DCT transform, and a frequency domain predictor 23b configured to perform residual encoding based on frequency domain prediction. In some implementations, the FD predictor is a FD predictor as disclosed in U.S. Provisional Patent Application No. 63 / 713,216 or European patent application number 24209441.5, which is herewith incorporated by reference in its entirety. The FD predictor 23 may be configured to apply a predictive coding operation to the FD frame wherein the predictive coding operation comprising sequentially processing the bands, to determine, for each of the bands, a predicted sample, a prediction residual, and a reconstructed sample. The prediction residual sample of each band forms the output from the FD predictor.
[0051] Similar to the architecture in figure 1, the encoder 23 is followed by a quantization and entropy coding block, and is preceded by a framing block 25. A coding adaptation block 26 is connected to receive information about the sampling frequencies, and resampling blocks 27 are configured to resample the input channels.
[0052] In this example, the transform 23a places restrictions on the processing frame length, which needs to be compatible with a computationally efficient transform length, e.g. 256, 512, 768, 1024, 2048, etc.
[0053] The resampling in blocks 27 is controlled by the coding adaptation block 26 in the same manner as in figure 1, based on the sampling frequencies and constraints on the processing time frame lengths.
[0054] Also, the framing in block 25 may be similar to the processing in figure 1. However, if the transform is a modified DCT (MDCT) then the framing block 25 provides overlapping time frames with 50% overlap. The coding adaptation block 26 also needs to communicate the selected processing frame length (and any overlap) to the transform 23a in each path 21, 22.
[0055] Finally, the coding adaptation block 26 also determines appropriate prediction orders for the frequency domain predictors 23b in each path 21, 22. Thes may be determined in various ways, for example they may be proportional to the processing frame length (i.e., a longer processing frame requires a higher prediction order).
[0056] Unless specifically stated otherwise, as apparent from the following discussions, it is appreciated that throughout the disclosure discussions utilizing terms such as “processing”, “computing”, “calculating”, “determining”, “analyzing” or the like, refer to the action and / or processes of a computer hardware or computing system, or similar electronic computing devices, that manipulate and / or transform data represented as physical, such as electronic, quantities into other data similarly represented as physical quantities.
[0057] It should be appreciated that in the above description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment. Thus, the claims following the Detailed Description are hereby expressly incorporated into this Detailed Description, with each claim standing on its own as a separate embodiment of this invention. Furthermore, while some embodiments described herein include some, but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the invention, and form different embodiments, as would be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0058] Furthermore, some of the embodiments are described herein as a method or combination of elements of a method that can be implemented by a processor of a computer system or by other means of carrying out the function. Thus, a processor with instructions for carrying out such a method or element of a method forms a means for carrying out the method or element of a method. Note that when the method includes several elements, e.g., several steps, no ordering of such elements is implied, unless specifically stated. Furthermore, an element described herein of an apparatus embodiment is an example of a means for carrying out the function performed by the element for the purpose of carrying out the embodiments of the invention. In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the invention may be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description.
[0059] The person skilled in the art realizes that the present invention by no means is limited to the preferred embodiments described above. On the contrary, many modifications and variations are possible within the scope of the appended claims.
[0060] Various aspects of the present disclosure may be appreciated from the following enumerated example embodiments (EEEs):
[0061] EEE 1. A method for encoding a multi-channel input signal, comprising: receiving at least two input channels, xn, each being acquired with a different sampling frequency, fsn, for each input channel, selecting an appropriate processing frame length, Ln, from a set of allowable processing frame lengths, based on the sampling frequency of each respective input channel, and encoding each input channel using the selected processing frame length.
[0062] EEE 2. The method according to EEE 1, further comprising: resampling at least one of the input channels, and selecting the appropriate processing frame lengths so that, for each pair of two input channels, a relationship between the sampling frequencies, after resampling, corresponds to a relationship between the selected processing frame lengths.
[0063] EEE 3. The method according to EEE 1, further comprising: resampling at least one of the input channels, for each channel, setting a multiplier defining a number of processing frames of the input channel to be included in a channel package, and selecting the appropriate processing frame lengths so that, for each pair of two input channels, a relationship between the sampling frequencies, after resampling, corresponds to a relationship between the channel packages.
[0064] EEE 4. The method according to EEE 2 or 3, wherein the resampling is up- sampling, so that the sampling frequency, after resampling, is greater than an original sampling frequency.
[0065] EEE 5. The method according to any one of the preceding EEEs, wherein the resampling is achieved by buffering the input channel in a buffer of appropriate length and sampling the buffer with a desired sampling frequency.
[0066] EEE 6. The method according to any one of the preceding EEEs, wherein the encoding involves a time-to-frequency transform, and wherein the set of allowable processing frame lengths are determined by transform lengths which are compatible with the time-to- frequency transform.
[0067] EEE 7. The method according to any one of the preceding EEEs, wherein the encoding is configured to operate on time domain frames having one of a set of predefined time domain frame lengths, and wherein the set of allowable processing frame lengths are determined by the set of time domain frame lengths.
[0068] EEE 8. The method according to any one of the preceding EEEs, wherein each processing frame length in the set of allowable processing frame lengths is a multiple of a power of two.
[0069] EEE 9. The method according to EEE 8, wherein the set of allowable processing frame lengths contains the processing frame lengths 2m, 2m+1, ... , 2n, where m and n are integers with n > m.
[0070] EEE 10. The method according to any one of the preceding EEEs, wherein the encoding involves frequency domain prediction and residual encoding.
[0071] EEE 11. The method according to EEE 10, wherein a prediction order in the frequency domain prediction of an input signal is selected based on the sampling frequency of the input signal.
[0072] EEE 12. The method according to any one of the preceding EEEs, further comprising transmitting a bitstream including the encoded signal.
[0073] EEE 13. The method according to EEE 12, when referring to EEE 2, wherein the bitstream also includes side information related to the resampling of the input signals.
[0074] EEE 14. The method according to any one of the preceding EEEs, wherein the input signal is a biomedical signal.
[0075] EEE 15. A method for decoding a multi-channel signal, comprising: receiving a bitstream, including an encoded multi-channel signal and side information; decoding the multi-channel signal; resampling at least one of the channels based on the side information.
[0076] EEE 16. A system for encoding a multi-channel input signal having at least two input channels, xn, each being acquired with a different sampling frequency, fsn, comprising: a coding adaptation block configured to, for each input signal, select an appropriate processing frame length, Ln, from a set of allowable processing frame lengths, based on the sampling frequency of each respective input channel, and a set of encoders, each configured to encode one of the input channel using the processing frame length selected for this channel.
[0077] EEE 17. The system according to EEE 16, further comprising a set of resampling blocks, each resampling block connected to receive one of the input channels, wherein the code adaptation block is configured to: control the set of resampling blocks to resample at least one of the input channels, and select the appropriate processing frame lengths so that, for each pair of input channels, a relationship between the sampling frequencies, after resampling, corresponds to a relationship between the selected processing frame lengths.
[0078] EEE 18. The method according to EEE 16, further comprising a set of resampling blocks, each resampling block connected to receive one of the input channels, wherein the code adaptation block is configured to: control the set of resampling blocks to resample at least one of the input channels, for each input channel, set a multiplier defining a number of processing frames of the input channel to be included in a channel package, and selecting the appropriate processing frame lengths so that, for each pair of two input channels, a relationship between the sampling frequencies, after resampling, corresponds to a relationship between the channel packages.
[0079] EEE 19. The method according to EEE 17 or 18, wherein the resampling is up- sampling, so that the sampling frequency, after resampling, is greater than an original sampling frequency.
[0080] EEE 20. The system according to one of EEEs 17 - 19, wherein each resampling block is configured to buffer an input channel in a buffer of appropriate length and to sample the buffer with a desired sampling frequency.
[0081] EEE 21. The system according to any one of EEEs 16 - 20, wherein each encoder involves a time-to-frequency transform, and wherein the set of allowable processing frame lengths are determined by transform lengths which are compatible with the time-to- frequency transform.
[0082] EEE 22. The method according to any one of EEEs 16 - 21, wherein each encoder is configured to operate on time domain frames having one of a set of predefined time domain frame lengths, and wherein the set of allowable processing frame lengths are determined by the set of time domain frame lengths.
[0083] EEE 23. The method according to any one of EEEs 16 - 22, wherein each processing frame length in the set of allowable processing frame lengths is a multiple of a power of two.
[0084] EEE 24. An encoder for encoding an input signal, the encoder comprising at least one processor configured to perform the method according to one of EEEs 1-15.
[0085] EEE 25. A computer program product comprising computer program code portions configured to perform the method according to one of EEEs 1-15 when executed on a computer processor.
[0086] EEE 26. A computer readable medium storing computer program code portions configured to perform the method according to one of EEEs 1-15 when executed on a computer processor.
Claims
CLAIMS1. A computer implemented method for encoding a multi-channel biomedical input signal, comprising: receiving at least two input channels, xn, each being acquired with a different sampling frequency, fsn, for each input channel, selecting a processing frame length, Ln, and encoding each input channel using the processing frame length, wherein the encoding of each input signal involves a time-to-frequency transform, wherein the processing frame length for each channel is selected from a set of available transform lengths which are compatible with the respective time-to-frequency transform, based on the sampling frequency of each respective input channel, in order to maintain time alignment of the input channels.
2. The method according to claim 1, further comprising: resampling at least one of the input channels, and selecting the appropriate processing frame lengths so that, for each pair of two input channels, a ratio between the sampling frequencies, after resampling, corresponds to a ratio between the selected processing frame lengths.
3. The method according to claim 1, further comprising: resampling at least one of the input channels, for each channel, setting a multiplier defining a number of processing frames of the input channel to be included in a channel package having a channel package length, and selecting the appropriate processing frame lengths so that, for each pair of two input channels, a ratio between the sampling frequencies, after resampling, corresponds to a ratio between the channel package lengths.
4. The method according to claim 2 or 3, wherein the resampling is up-sampling, so that the sampling frequency, after resampling, is greater than an original sampling frequency.
5. The method according to any one of the preceding claims, wherein the resampling is achieved by buffering the input channel in a buffer of appropriate length and sampling the buffer with a desired sampling frequency.
6. The method according to any one of the preceding claims, wherein the encoding is configured to operate on time domain frames having one of a set of predefined time domain frame lengths, and wherein the set of available processing frame lengths are determined by the set of time domain frame lengths.
7. The method according to any one of the preceding claims, wherein each processing frame length in the set of available processing frame lengths is a multiple of a power of two.
8. The method according to claim 7, wherein the set of allowable processing frame lengths contains the processing frame lengths 2m, 2m+1, ... , 2n, where m and n are integers with n > m.
9. The method according to claim 8, wherein the time-to-frequency transform is MDCT, and the set of available processing frame lengths include 512, 768 and 1024.
10. The method according to any one of the preceding claims, wherein the encoding involves frequency domain prediction and residual encoding.
11. The method according to claim 10, wherein a prediction order in the frequency domain prediction of an input signal is selected based on the sampling frequency of the input signal.
12. The method according to any one of the preceding claims, further comprising transmitting a bitstream including the encoded signal.
13. The method according to claim 12, when referring to claim 2, wherein the bitstream also includes side information related to the resampling of the input signals.
14. A method for decoding a multi-channel signal, comprising: receiving a bitstream, including an encoded multi-channel signal and side information; decoding the multi-channel signal; resampling at least one of the channels based on the side information.
15. A system for encoding a multi-channel biomedical input signal having at least two input channels, xn, each being acquired with a different sampling frequency, fsn, comprising: a coding adaptation block configured to, for each input signal, select a processing frame length, Ln, and a set of encoders, each configured to encode one of the input channel using the processing frame length, wherein the encoders are configured to apply a time-to-frequency transform to each input signal, wherein the coding adaptation block is configured to select the processing frame length for each channel from a set of available transform lengths which are compatible with the respective time-to-frequency transform, based on the sampling frequency of each respective input channel, in order to maintain time alignment of the input channels.
16. The system according to claim 15, further comprising a set of resampling blocks, each resampling block connected to receive one of the input channels, wherein the code adaptation block is configured to: control the set of resampling blocks to resample at least one of the input channels, and select the appropriate processing frame lengths so that, for each pair of input channels, a ratio between the sampling frequencies, after resampling, corresponds to a ratio between the selected processing frame lengths.
17. The method according to claim 15, further comprising a set of resampling blocks, each resampling block connected to receive one of the input channels, wherein the code adaptation block is configured to: control the set of resampling blocks to resample at least one of the input channels, for each input channel, set a multiplier defining a number of processing frames of the input channel to be included in a channel package having a channel package length, and selecting the appropriate processing frame lengths so that, for each pair of two input channels, a ratio between the sampling frequencies, after resampling, corresponds to a ratio between the channel package lengths.
18. The method according to claim 16 or 17, wherein the resampling is up-sampling, so that the sampling frequency, after resampling, is greater than an original sampling frequency.- 18 -19. The system according to one of claims 16 - 18, wherein each resampling block is configured to buffer an input channel in a buffer of appropriate length and to sample the buffer with a desired sampling frequency.
20. The method according to any one of claims 15 - 19, wherein each encoder is configured to operate on time domain frames having one of a set of predefined time domain frame lengths, and wherein the set of allowable processing frame lengths are determined by the set of time domain frame lengths.
21. The method according to any one of claims 15 - 20, wherein each processing frame length in the set of allowable processing frame lengths is a multiple of a power of two.
22. An encoder for encoding an input signal, the encoder comprising at least one processor configured to perform the method according to one of claims 1-13.
23. A computer program product comprising computer program code portions configured to perform the method according to one of claims 1-13 when executed on a computer processor.
24. A non-transitory computer readable medium storing computer program code portions configured to perform the method according to one of claims 1-13 when executed on a computer processor.- 19 -
Citation Information
Patent Citations
Scalable compressed audio bit stream and codec using a hierarchical filterbank and multichannel joint coding
US20070063877A1
Method, apparatus and systems for audio decoding and encoding
US20230274755A1
EP24209441A
US63713216P