Multi-level vector quantization for audio coding
By adopting segmented structure and DCT transform optimized multi-level vector quantization technology in audio coding, the problems of high storage space and search complexity are solved, and low-storage and low-complexity audio coding and decoding are achieved.
Patent Information
- Application Number
- CN202480014456.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-23
- Filing Date
- 2024-02-14
- Publication Date
- 2025-10-24
AI Technical Summary
Existing multi-level vector quantization technology has problems of high storage space and search complexity in audio coding, especially the first-level quantization requires a large amount of storage space and high-complexity operations.
A segmented codebook is adopted to reduce storage space and search complexity through offline optimization and DCT transformation. The segmented truncation length and optimized candidate update mechanism are used in combination with inverse DCT transformation to reconstruct the vector.
It effectively reduces storage space and search complexity, realizes multi-level vector quantization with low storage and low operation complexity, and is suitable for audio encoders and decoders.
Smart Images

Figure CN120836056A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is a continuation-of-U.S. patent application serial number 63 / 454,166 filed on March 23, 2023, which claims the benefit of U.S. provisional patent application serial number 63 / 448,440 filed on February 27, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure relates generally to encoding and decoding, and more particularly to a multi-level vector quantization method and related encoders and / or decoders supporting multi-level vector quantization. Background Art
[0004] Vector quantization (VQ) has been widely used in the development of data compression technology. Notably, VQ is a highly efficient data compression technique that constructs multiple scalar data columns into vectors and performs overall quantization in the vector space. As a result, data is compressed without losing much information.
[0005] A general random codebook of 24 coefficients / vectors with 37-bit dimension would require: 2 37 *24=3*2 40 =3.29*10 12 In most real-world implementations, it is impractical to design, search, and store such a codebook.
[0006] In the 1980s and 1990s, the concepts of split-VQ and multi-level VQ (MSVQ) were developed to achieve more reasonable storage and search complexity while still maintaining considerable vector quantization gains. In the training of these VQs, the coefficient correlation characteristics are still appropriately exploited.
[0007] An example split VQ is the adaptive multi-rate wideband (AMR-WB) instantaneous spectral frequency (ISF) VQ described in the 3rd Generation Partnership Project (3GPP) technical standard (TS) 26.190 in the section "5.2.5 Quantization of ISP coefficients". Therein the total number of coefficients is 16, which is split in a first stage into 9 coefficients (8 bits) and 7 coefficients (8 bits). Also, in a second stage, the VQ is further split into {3,3,3} coefficients with (6,7,7) bits and a higher stage 2 split of {3,4} using (5,5) bits. In total 8+8+6+7+7+5+5 = 46 bits are used and the total number of ROM table entries in Matlab syntax becomes: sum((2.^[8 8]).*[9 7])+sum((2.^[6 7 7]).*[3 3 3])+sum((2.^[5 5]).*[3 4]) = 5.2kW (Word 16), i.e. a much more realistic table ROM number.
[0008] A less complex alternative to trained random (SplitVQ, MSVQ) codebooks are codebooks with a given algebraic or lattice structure, which are easier to search, but which will have suboptimal Voronoi regions and are typically only efficient for certain input vector distributions (e.g. Gaussian, Laplacian, etc.). Furthermore, efficient indexing schemes (e.g. constructing the final codebook vector from the bitstream) can become expensive, and further, lattice structures are typically not very flexible in terms of vector length. Examples of lattice quantizers are the D8 lattice, the RE8 lattice and the pyramid VQ (PVQ). Spherical PVQ is also used in enhanced voice services (EVS) and Internet Engineering Task Force (IETF) Opus for encoding Gaussian sources.
[0009] Another possibility is to transform the input signal to a better domain for efficient and fast quantization. For example, a signal optimized Karhunen-Love transform (KLT) or a more general discrete cosine transform (DCT) can be used to obtain energy (i.e. via KLT) or frequency (i.e. via DCT) relations for the input signal. SUMMARY
[0010] There are certain challenges. Current solutions for the first stage (and subsequent stages) of MSVQ typically require large storage space. For example, in the case of EVS, the frequency domain comfort noise generation (FD-CNG) VQ in the first stage uses 128 levels (7 bits) with 24 coefficients per level, resulting in (128 x 24) = 3072 words of ROM storage (i.e., typically a word can be a 16-bit Wordl6 integer, or a word can be a single-precision 32-bit floating point number). In the EVS-FD-CNG-VQ (Enhanced Voice Services - Frequency Domain Comfort Noise Generation - Vector Quantization) implementation, this corresponds to 6 kilobytes (kB) with 8 bits per byte.
[0011] The entire EVS 37-bit (7+5*6) bit FD-CNG VQ (all stages) is using (128+5x64)*24 = 10752 Wordl6 or 20 kB of ROM. This is twice the ROM size of the 46-bit AMR-WB ISF-VQ.
[0012] For the EVS FD-CNG-VQ, the unstructured first stage is using about 30% of the quantizer's ROM space while also using 19% (7 bits / 37 bits) of the information space.
[0013] The lack of an efficient structure in the first stage makes it difficult to provide a low storage space and low complexity search solution.
[0014] The current per-vector iteration update of Nc (i.e., the number of candidates to be kept for the 2nd stage) has a very high worst case "weighted millions of operations per second" (WMOPS) complexity because the active list of Nc (8 in EVS) is potentially updated for every analyzed codebook vector, i.e., the Nc length candidate list can need to be updated 128 times.
[0015] Certain aspects of the present disclosure and embodiments thereof can provide solutions to these or other challenges. A transformed segment structure with different truncation lengths (e.g., corresponding to different levels of detail) in each segment is introduced to the first stage codebook.
[0016] An optimized inner loop keeps and / or maintains a pair of best candidates (e.g., two lists) for each segment, effectively searching each segment with its individual truncation length.
[0017] The number of segments Ns in the first stage is kept around Ns ~ = Nc / 2. Thus, at the end of the search, around Nc = 2*Ns candidates will be available on demand.
[0018] To further improve the quality of the remaining Nc candidates, the list of Nc candidates is updated by using the stored circular list of the most recent first stage vectors. This update is performed based on evaluating the first and second best candidate neighbors (e.g., in terms of mean square error, MSE) against the worst existing entry in the original list of Nc candidates.
[0019] To further improve the storage and retrieval characteristics of the first stage, the actual coefficients of the segmented vectors are stored in the Word8 quantization domain and special scaling (e.g., implemented as a shift factor) is performed on each column of the segment.
[0020] According to some embodiments, the disclosed subject matter includes a method implemented by an encoder, the method comprising obtaining a discrete cosine transform (DCT) target vector. The method further comprises implementing a suboptimal pairwise inner search in each segment of a codebook having a plurality of segments to determine a pairwise initial candidate set from each segment of the plurality of segments, forming a plurality of pairwise initial candidates from the suboptimal pairwise inner search, wherein each segment has a truncation vector that is different from the truncation vectors of other segments of the plurality of segments. The method comprises implementing a post-optimization on the plurality of pairwise initial candidates to replace ones of the plurality of pairwise initial candidates using a candidate neighbor list, forming a plurality of final candidates. The method comprises reconstructing the plurality of final candidates using an inverse type-II discrete cosine transform (DCT-II) to transform the plurality of final candidates into a plurality of reconstructed final candidates in an original domain. The method comprises providing the plurality of reconstructed final candidates to a second stage of a multi-stage vector quantizer.
[0021] Certain embodiments can provide one or more of the following technical advantages. The introduced codebook structure (e.g., transformed and selectively truncated into Ns segments) allows for fast search using less operations (e.g., WMOPS).
[0022] Pairwise candidate search within each codebook segment allows for lower complexity (WMOPS) inner loop of the first stage. By efficiently swapping in and out new best entries in the pair (e.g., two lists).
[0023] The codebook vector is obtained from memory (e.g., ROM) using segment- and coefficient-column-specific scaling (stored as exponent of shift factor). The codebook vector mantissa values (e.g., “segment codebook vector” values) are represented by signed bytes (Word8) respectively, thus additionally enabling fast ROM read access as a pair of bytes (Word16) (e.g., efficient Word8 mantissa retrieval in a digital signal processor (DSP) in case of equal vector coefficient truncation length for all segments), and furthermore, the first stage codebook search of a segment allows parallelization of the first stage (e.g., stage #1) VQ inner loop. The Word8 and shift factor representation for each coefficient means that the coefficients are selectively truncated within the dynamic range of each column. The truncation for each segment means that each vector is also truncated at the level of detail. Using DCT, the truncation corresponds to removing a number of high frequency DCT components. As used herein, the term “Word8” can refer to a signed 8-bit integer in ITU-T G.191 basic operators. Likewise, the term “Word16” can refer to a signed 16-bit integer in ITU-T G.191 basic operators.
[0024] According to some other embodiments, a method in a decoder for reconstructing a target vector includes receiving an index containing a plurality of segments, corresponding shift values, and a global offset value vector. The method includes performing a segment- and coefficient-wise upshift of vectors in the plurality of segments using column shift (e.g., “col_shift”) values to form upshifted vectors within each of the plurality of segments. The method includes performing an inverse discrete cosine transform, DCT Type II transform, on the upshifted segments. The method includes downscaling the output vectors from the DCT Type II transform to original unscaled FDCNG domain vectors. The method includes adding the global offset value vector back to the unscaled frequency domain comfort noise generation, FD-CNG, domain vectors to form the target vector. BRIEF DESCRIPTION OF DRAWINGS
[0025] The accompanying drawings, which are included to provide a further understanding of the disclosure and are incorporated in and constitute a part of this application, illustrate certain non-limiting embodiments of the inventive concepts. In the drawings:
[0026] Figure 1 is a block diagram of a system view of FD-CNG-CQ with low ROM and low WMOPS according to some embodiments;
[0027] Figure 2 is a diagram of an example of an operating environment of an encoder according to some embodiments;
[0028] Figure 3 is a block diagram of an audio encoder according to some embodiments,
[0029] Figure 4 is a block diagram of an audio decoder according to some embodiments;
[0030] Figure 5 is a block diagram of a host according to some embodiments;
[0031] Figure 6 is a block diagram of an exemplary virtualized environment in which an encoder or decoder or some components of an encoder or decoder can be implemented according to some embodiments;
[0032] Figure 7 is a block diagram of a first stage FD-CNG VQ search according to some embodiments;
[0033] Figure 8 is a block diagram of a system view of FD-CNG-CQ with low ROM and low WMOPS supporting shorter input target vectors according to some embodiments;
[0034] Figure 9 is a block diagram of a first stage FD-CNG VQ search supporting shorter input target vectors according to some embodiments;
[0035] Figure 10 is a diagram showing the operation of an encoder implementing DCT BASIS vector extrapolation to extend short input target vectors according to some embodiments.
[0036] Figure 11 is a block diagram showing FD-CNG vector reconstruction for MSVQ first stage according to some embodiments;
[0037] Figures 12 to 18 is a flow diagram showing the operation of an encoder according to some embodiments; and
[0038] Figure 19 is a flow diagram showing the operation of a decoder according to some embodiments. DETAILED DESCRIPTION
[0039] Some embodiments as contemplated herein will now be described more fully with reference to the accompanying drawings. Examples are provided by way of example only and are not intended to limit the scope of what can be claimed. The present concepts can be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and fully convey the scope of the present concepts to those skilled in the art. It should also be noted that these embodiments are not mutually exclusive. Components in one embodiment can be assumed to exist / use in another embodiment by default.
[0040] As previously mentioned, current solutions for MS-VQ first stage (and subsequent stages) typically still require large storage space. For example, in the case of EVS for FD-CNG VQ, the first stage uses 128 levels (7 bits), and 24 coefficients per level, resulting in (128 x 24) = 3072 words (Word 16) of ROM storage. In the EVS FD-CNG VQ implementation, this corresponds to 6k bytes, where each byte is 8 bits. The entire EVS 37-bit (7+5*6) FD-CNG VQ (all stages) is using (128+5x64)*24 = 10752 Word 16 or 20k bytes of ROM. This is twice the ROM size of the 46-bit AMR-WB ISF-VQ.
[0041] In some embodiments, the systems described herein show how to reduce storage space, which includes the first stage of a multi-stage VQ, to improve the table ROM and worst case WMOPS characteristics of the 3GPP EVS-26.445 FD-CNG VQ spectral envelope quantization.
[0042] For example, the EVS FD-CNG VQ is used to quantize the spectral envelope used in the Silence Insertion Descriptor (SID) frame to 37 bits. The decoded SID frame can then be used to generate a time-domain background signal.
[0043] However, the first stage (i.e., stage #1) optimization of the MSVQ as described below can also be applied to any audio codec vector parameter that uses multi-stage vector quantization. For example, the spectral envelope quantized by the MSVQ can be used for spectral noise shaping (SNS) during active music segments, or can be used to quantize and represent the spectral envelope used to synthesize active unvoiced speech segments.
[0044] Figure 1 An example enhanced structured 1 -stage search (i.e., first stage search) system 100 is shown in FIG. 1. The structural aspects are shown as being performed a priori offline, and the search and structural aspects employed during the Nc candidate vector determination.
[0045] Turning to Figure 1, the initial 1-level VQ codebook 102 trained by LBG (or Kmeans) has the optimal Voronoi region, as described above, using 128 levels (7 levels) and 24 coefficients per level, resulting in a ROM storage of (128x24)=3072 words (Word16). Codebook 102 enables a complete search of 8 candidates (i.e., candidate vectors that can be indicated and / or pointed to by a codebook index). According to various embodiments, codebook 102 is converted by DCT24 into a DCT24 truncated (8, 10, 16, 18) table Word8 codebook 104 (e.g., see codebook 104 with truncated segments 105 including vectors of length 8 coefficients), where the number of segments Ns is 4 segments and the maximum number of coefficients is 18, and enables a complete search of up to 8 candidates. As shown in Table 1 below, the truncated segment 105 contains 16 vectors (eg, see nSeg[0] == 16), each vector being 8 coefficients long, for a total of 128 values.
[0046] Codebook 104 is created offline, resulting in a ROM storage of 987 words (Word16), which is approximately 30% of the size of codebook 102. In some embodiments, codebook 104 has 4 segments (i.e., the number of segments Ns=4). Optimization of codebook 104 (e.g., Figure 1 8. The codebook 104 is also implemented offline to create vectors in the codebook 104 (as shown in block 108 in FIG. 1 ). This may include performing an inverse DCT (IDCT) transform on the vectors in the codebook 104 to optimize the vectors in the codebook 104, as shown by the IDCT block 106 and the optimization block 108. Notably, spectral distortion (SD) may be maintained when creating the codebook 104. The optimization also includes creating a "nearest neighbor" index circular list 110, which contains vectors pointing to the next index or the previous index in approximate mean square error (MSE) neighbor order. In some embodiments, the nearest neighbor index circular list 110 has 128 entries in ROM word 8.
[0047] like Figure 1 Blocks 112-120 are shown as part of a FD-CNG VQ search according to various embodiments, which feeds and / or provides the eight best candidates to the second stage (e.g., stage 2) of the multi-stage VQ, as shown in block 122, where further processing of the VQ is performed. Before describing these blocks, the parameters used should first be defined in terms of parameter name, dimension, value, and brief definition, as shown in Table 1 below.
[0048]
[0049]
[0050]
[0051] Table 1
[0052] One feature shared by various embodiments of the disclosed subject matter is that there is a performance saving when comparing the original table ROM space of 3072 single precision floating point (32 bits per coefficient) yielding a total ROM storage of 12288 bytes in the EVS floating point specification to the first stage of the FD-CNG VQ (first stage uses 128 levels (7 bits) with 24 coefficients per level). These performance savings are indicated in Table 2. It should be noted that the table ROM space in the corresponding EVS fixed point specification is 3072 16-bit integers (Word 16's) resulting in a total ROM storage of 6144 bytes.
[0053]
[0054]
[0055] Table 2
[0056] In some embodiments, the first stage codebook 104 is structured such that it contains segments (e.g., 4) with different DCT truncation lengths (e.g., 8, 10, 16, 18). Each segment has a fixed common truncation length (e.g., number of remaining coefficients) for its vector set. For example, in the context of the present disclosure, a length of 24 coefficients would indicate that the vector is of length without truncation, while a truncation length of 8 coefficients would indicate that the vector has been truncated and / or reduced to a length of 8 coefficients. The resulting dimensionality reduction reduces both storage size and search complexity.
[0057] In the present disclosure, a design with Ns = 4 segments (with different truncation lengths) will be used. Each segment has different frequency characteristics, including a range from low frequency codebooks to mid frequency codebooks to high frequency codebooks.
[0058] In some embodiments, the truncation lengths have been optimized offline. One example simple design method to establish truncation lengths is to simply set relative energy requirements as follows:
[0059] In a first step, a non-truncated random basis codebook of vector length N FDCNG is trained using a kmeans or LBG VQ training algorithm.
[0060] For each vector included in the basis codebook 102, an initial individual truncation length is set to the length where the vast majority (e.g., 95%) of the remaining vector energy is.
[0061] In a second step, the initial vector length is increased individually such that Ns subsets with equal truncated length can be created. Then, each of these subsets can be represented as segments with equal truncated length and all subsets have at least 95% of the remaining energy.
[0062] In some embodiments, additional criteria can be added such that the reconstructed point of each vector in the truncated codebook vector set should not move to the index Voronoi region of another vector in the original kmeans / LBG trained basis codebook (CB). In further embodiments, the allowed movement away from the original reconstructed point can be limited so as not to deviate too far from the original basis codebook’s Voronoi region setup.
[0063] Before describing the first stage FD-CNG VQ search and reconstruction, an overview of the operating environment, encoder, decoder, host, and virtualization environment will first be described herein.
[0064] Figure 2 A block diagram of an example operating environment 200 in which various embodiments of the present disclosure can be implemented is shown. Turning to Figure 2 In the example operating environment 200, an encoder 202 receives data to be encoded, such as an audio file, from an entity, such as a host 206, and / or from a memory 208 over a network 204. In various embodiments, the encoder 202 is a parametric stereo encoder. In some embodiments, the host 206 can communicate directly with the encoder 202. In some embodiments, the encoder 202 can encode the audio file as described herein and store the encoded audio file in the memory 208 or send the encoded video file to a decoder 212 via a network 210. In various embodiments, the decoder 212 is a parametric stereo decoder. The decoder 212 decodes the audio file and sends the decoded audio file to an audio player 214 for playback. The audio player 214 can be or be included in a user device, terminal, mobile phone, etc. In other embodiments, the host 206 can send the encoded audio file to the decoder 212 via the network 210.
[0065] Figure 3An audio encoder 202 according to some embodiments is shown, where the audio encoder 202 is implemented as a standalone device. As used herein, an audio encoder refers to a device that is capable of, configured, arranged and / or operable to encode objects and communicate with network nodes, encoders and / or decoders. Examples of audio encoders include, but are not limited to, a smartphone, a mobile phone, a cellular phone, a Voice over IP (VoIP) phone, a wireless local loop phone, a desktop computer, a Personal Digital Assistant (PDA), a wireless camera, a gaming console or device, a storage device, a playback device, a wearable terminal device, a wireless endpoint, a mobile station, a tablet, a laptop computer, a laptop computer embedded equipment (LEE), a laptop computer mounted equipment (LME), a smart device, a wireless customer premises equipment (CPE), a vehicle-mounted or embedded / integrated wireless device, etc.
[0066] The audio encoder can support device-to-device (D2D) communication, e.g., by implementing 3GPP standards for sidelink communication, dedicated short-range communications (DSRC), vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), or vehicle-to-anything (V2X). In other examples, the encoder can not necessarily have a user in the sense of a human user that owns and / or operates the related device.
[0067] The audio encoder 202 includes processing circuitry 302 that is operatively coupled to an input / output interface 306, a power source 308, a memory 310, a communication interface 312, and / or any other component(s) or combination of Figure 3 components as illustrated in FIG. 3. The level of integration of the components can vary from one design to another, and from one decoder to another. Moreover, a certain decoder can not include one or more of the components of FIG. 3. Additionally, a certain decoder can include one or more other components not expressly shown in FIG. 3.
[0068] The processing circuitry 302 is configured to process instructions and data, and can be configured to implement any sequential state machine operative to execute instructions stored in the memory 310, such as a machine-readable computer program, as an ordered listing of executable instructions. The processing circuitry 302 can be implemented in one or more hardware-implemented state machines, for example in discrete logic, field programmable gate array (FPGA), application- specific integrated circuit (ASIC), etc.; in programmable logic and appropriate firmware; in one or more stored computer programs, general purpose processors, such as a microprocessor or digital signal processor (DSP), and appropriate software; or any combination thereof. For example, the processing circuitry 302 can include multiple central processing units (CPUs).
[0069] In this example, the input / output interface 306 can be configured to provide one or more interfaces to input devices, output devices, or one or more input and / or output devices. Examples of output devices include a speaker, a sound card, a video card, a display, a monitor, a printer, an actuator, an emitter, a smartcard, another output device, or any combination thereof. The input devices can allow a user to capture information into the audio encoder 202. Examples of input devices include a touch-sensitive or presence-sensitive display, a camera (e.g., a digital camera, a digital video camera, a web camera, etc.), a microphone, a sensor, a mouse, a trackball, a directional pad, a trackpad, a scroll wheel, a smartcard, etc. The presence-sensitive display can include a capacitive or resistive touch sensor to sense inputs from a user. The sensor can be, for example, an accelerometer, a gyroscope, a tilt sensor, a force sensor, a magnetometer, an optical sensor, a proximity sensor, a biometric sensor, etc., or any combination thereof. The output devices can use the same types of interfaces as the input devices. For example, a universal serial bus (USB) port can be used to provide input to both input and output devices.
[0070] In some embodiments, the power supply 308 is configured as a battery or battery pack. Other types of power supplies, such as an external power supply (e.g., an electricity outlet), a photovoltaic device, or a capacitor, can be used. The power supply 308 can also include power circuitry to deliver power from the power supply 308 and / or an external power source, e.g., via an input circuit or interface such as an electrical cable, to the various components of the audio encoder 202. Delivering power can be used, for example, to charge the power supply 308. The power circuitry can perform any formatting, converting, or other modification to the power from the power supply 308 to make the power suitable for use by the respective components of the audio encoder 202 being powered.
[0071] The storage 310 can be or include memory such as random access memory (RAM), read only memory (ROM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), electronically erasable programmable read only memory (EEPROM), a magnetic disk, an optical disk, a hard disk, a floppy disk, a file, a tape, a flash drive, etc. In one example, the storage 310 includes one or more applications 314, such as an operating system, a web browser application, a widget or widget engine, or other applications, and corresponding data 316. The storage 310 can store any of a variety of different operating systems, or combinations of operating systems, for use by the audio encoder 202.
[0072] The memory 310 can be configured to include a plurality of physical drive units, such as independent disk drives, flash memory, USB flash drives, external hard drives, thumb drives, pen drives, key drives, High-Density Digital Versatile Disc (DVD) optical drives, internal hard drives, Blu-Ray optical drives, Holographic Digital Data Storage (HDDS) optical drives, external mini-dual in-line memory modules (DIMMs), synchronous dynamic random access memory (SDRAM), external micro-DIMMs, smart cards, such as Universal Integrated Circuit Cards (UICCs) in the form of Subscriber Identity Modules (SIM), Universal SIM (USIM), and / or ISIM, other memory, or any combination thereof. The UICC may, for example, be an embedded UICC (eUICC), an integrated UICC (iUICC), or a removable UICC commonly referred to as a "SIM card." The memory 310 can allow the audio encoder 202 to access instructions, application programs, or data stored on transitory or non-transitory storage media to offload or upload data. An article of manufacture, such as one utilizing a communication system, can have tangibly embodied therein or stored in storage thereon the memory 310, which can be or include a device readable storage medium.
[0073] The processing circuitry 302 can be configured to communicate with an access network or other networks using the communication interface 312. The communication interface 312 can include one or more communication subsystems and can include or be communicably coupled to an antenna 322. The communication interface 312 can include one or more transceivers used for communicating with one or more remote transceivers of another device capable of wireless communication (e.g., a UE or a network node of an access network), for example. Each transceiver can include a transmitter 318 and / or a receiver 320 adapted to provide network communications (e.g., optical, electrical, frequency allocation, etc.). Moreover, the transmitter 318 and the receiver 320 can be coupled to one or more antennas, such as the antenna 322, and can share circuit components, software, or firmware, or alternatively can be implemented separately.
[0074] In the illustrated embodiment, the communication functionality of the communication interface 312 can include cellular communication, Wi-Fi communication, LPWAN communication, data communication, voice communication, multimedia communication, short-range communication such as Bluetooth, near-field communication, location-based communication such as using the Global Positioning System (GPS) to determine a location, another similar communication functionality, or any combination thereof. The communication can be implemented according to one or more communication protocols and / or standards, e.g., IEEE 802.11, Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), GSM, LTE, New Radio (NR), UMTS, WiMax, Ethernet, Transmission Control Protocol / Internet Protocol (TCP / IP), Synchronous Optical Networking (SONET), Asynchronous Transfer Mode (ATM), QUIC, Hypertext Transfer Protocol (HTTP), etc.
[0075] Regardless of the type of sensor, the audio encoder can provide an output of encoded data via a wireless connection to a network node through its communication interface 312.
[0076] When in the form of an Internet of Things (IoT) device, the audio encoder can be a device for one or more application areas, including but not limited to: urban wearable technology, extended industrial applications, and healthcare. Non-limiting examples of such IoT devices are devices that are or are embedded in: a connected refrigerator or freezer, a television, a connected lighting device, an electricity meter, a robotic vacuum cleaner, a voice-controlled smart speaker, a home security camera, a thermostat, an electric door lock, a connected doorbell, an autonomous vehicle, a surveillance system, a weather monitoring device, a parking monitoring device, an electric vehicle charging station, a smart watch, a fitness tracker, a head-mounted display for augmented reality (AR) or virtual reality (VR), a wearable device for haptic or sensory augmentation. The decoder in the form of an IoT device includes, in addition to the other components described in connection with the audio encoder 202 shown in Figure 3
[0077] Figure 4 An audio decoder 212 (e.g., parametric stereo encoder) according to some embodiments is shown, where the audio decoder 212 is implemented as a standalone device. As used herein, an audio decoder refers to a device that is capable of, configured, arranged, and / or operable to decode objects and communicate with network nodes, encoders, and / or decoders. Examples of audio decoders include, but are not limited to, a smartphone, a mobile phone, a cellular phone, a Voice over IP (VoIP) phone, a wireless local loop phone, a desktop computer, a personal digital assistant (PDA), a wireless camera, a gaming console or device, a storage device, a playback device, a wearable terminal device, a wireless endpoint, a mobile station, a tablet, a laptop computer, a laptop computer embedded equipment (LEE), a laptop computer mounted equipment (LME), a smart device, a wireless customer premises equipment (CPE), an in-vehicle or embedded / integrated wireless device, etc.
[0078] An audio decoder can support device-to-device (D2D) communication, for example, by implementing 3GPP standards for sidelink communication, dedicated short-range communications (DSRC), vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), or vehicle-to-anything (V2X). In other examples, a decoder can not necessarily have a user in the sense of a human user that owns and / or operates the related device.
[0079] The audio decoder 212 includes processing circuitry 402 that is operatively coupled to an input / output interface 406, a power source 408, a memory 410, a communication interface 412, and / or any other component(s) or combination of Figure 4 all or a subset of the components shown in FIG. 4. The level of integration can vary from device to device. Furthermore, a certain decoder can contain multiple instances of a component, such as multiple processors, memories, transceivers, transmitters, receivers, etc.
[0080] The processing circuitry 402 is configured to process instructions and data, and can be configured to implement any sequential state machine operative to
[0081] In this example, the input / output interface 406 can be configured to provide one or more interfaces to an input device, output device, or one or more input and / or output devices. Examples of output devices include a speaker, sound card, video card, display, monitor, printer, actuator, transmitter, smartcard, another output device, or any combination thereof. The input devices can allow a user to capture information into the audio decoder 212. Examples of input devices include a touch-sensitive or presence-sensitive display, a camera (e.g., a digital camera, a digital video camera, a web camera, etc.), a microphone, a sensor, a mouse, a trackball, a directional pad, a trackpad, a scroll wheel, a smartcard, etc. The presence-sensitive display can include a capacitive or resistive touch sensor to sense input from a user. The sensor can be, for example, an accelerometer, a gyroscope, a tilt sensor, a force sensor, a magnetometer, an optical sensor, a proximity sensor, a biometric sensor, etc., or any combination thereof. The output devices can use the same types of interfaces as the input devices. For example, a universal serial bus (USB) port can be used to provide input to both input and output devices.
[0082] In some embodiments, the power supply 408 is configured as a battery or battery pack. Other types of power supplies, such as an external power supply (e.g., an electricity outlet), a photovoltaic device, or a capacitor, can be used. The power supply 408 can also include a power supply circuit to deliver power from the power supply 408 and / or an external power source, via an input circuit or interface such as an electrical cable, to various components of the audio decoder 212. Delivering power can be used, for example, to charge the power supply 408. The power supply circuit can perform any formatting, conversion, or other modification to the power from the power supply 408 to make the power suitable for use by the respective components of the audio decoder 212 being powered.
[0083] The storage 410 can be or include memory such as random access memory (RAM), read only memory (ROM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), electronically erasable programmable read only memory (EEPROM), magnetic disks, optical disks, hard drives, floppy disks, flash memory devices, or any other suitable device, however. In one example, the storage 410 includes one or more applications 414, such as an operating system, a web browser application, a widget or gadget engine, or other applications, and corresponding data 416. The storage 410 can store any of a variety of different operating systems, or combinations of operating systems, for use by the audio decoder 212.
[0084] Memory 410 can be configured to include a plurality of physical drive units, such as Redundant Array of Independent Disks (RAID), flash memory, USB flash drives, external hard drives, thumb drives, pen drives, key drives, High-Density Digital Versatile Disc (HD-DVD) optical disc drives, internal hard disk drives, Blu-Ray optical disc drives, Holographic Digital Data Storage (HDDS) optical disc drives, external mini-dual in-line memory modules (DIMMs), synchronous dynamic random access memories (SDRAM), external micro-DIMMs, smart cards, such as Universal Integrated Circuit Cards (UICCs) in the form of Subscriber Identity Modules (SIM), Universal Serial Bus (USB) flash drives, other memory, or any combination thereof. UICC can be, for example, an embedded UICC (eUICC), an integrated UICC (iUICC), or a removable UICC commonly referred to as a "SIM card." Memory 310 can allow audio decoder 212 to access instructions, application programs, or data stored on transitory or non-transitory memory media to offload or upload data. An article of manufacture, such as one utilizing a communication system, can have tangibly embodied therein or stored in storage 410, which can be device readable storage medium.
[0085] Processing circuitry 402 can be configured to communicate with an access network or other networks using communication interface 412. Communication interface 412 can include one or more communication subsystems and can include or be communicably coupled to one or more antennas 422. Communication interface 412 can include one or more transceivers used to communicate with one or more remote transceivers of another device capable of wireless communication such as a UE or a network node of an access network, for example. Each transceiver can include transmitter 418 and / or receiver 420 adapted to provide network communications (e.g., optical, electrical, frequency allocations, etc.). Additionally, transmitter 418 and receiver 420 can be coupled to one or more antennas, such as antenna 422, and can share circuit components, software, or firmware, or alternatively can be implemented separately.
[0086] In the illustrated embodiment, the communication functionality of the communication interface 412 can include cellular communication, Wi-Fi communication, LPWAN communication, data communication, voice communication, multimedia communication, short-range communication such as Bluetooth, near-field communication, location-based communication such as using the global positioning system (GPS) to determine a location, another similar communication functionality, or any combination thereof. The communication can be implemented in accordance with one or more communication protocols and / or standards, such as IEEE 802.11, code division multiple access (CDMA), wideband code division multiple access (WCDMA), GSM, LTE, New Radio (NR), UMTS, WiMax, Ethernet, transmission control protocol / Internet protocol (TCP / IP), synchronous optical networking (SONET), asynchronous transfer mode (ATM), QUIC, hypertext transfer protocol (HTTP), and / or the like.
[0087] Regardless of the type of sensor, the audio decoder can provide an output of the decoded data via a wireless connection to a network node through its communication interface 412.
[0088] When in the form of an Internet of Things (IoT) device, the audio decoder can be a device for one or more application areas, including but not limited to: urban wearable technology, extended industrial applications, and healthcare. A non-limiting example of such an IoT device is a device that is or is embedded in: a connected refrigerator or freezer, a television, a connected lighting device, an electricity meter, a robotic vacuum cleaner, a voice-controlled smart speaker, a home security camera, a thermostat, an electric door lock, a connected doorbell, an autonomous vehicle, a surveillance system, a weather monitoring device, a parking monitoring device, an electric vehicle charging station, a smart watch, a fitness tracker, an augmented reality (AR) or virtual reality (VR) head-mounted display, a wearable device for haptic or sensory augmentation. The decoder in the form of an IoT device can include, in addition to the other components described in connection with the audio decoder 212 shown in Figure 4 Fig. 1, circuitry and / or software that is dependent on the intended application of the IoT device.
[0089] Figure 5 is a block diagram of a host 206 in accordance with various aspects described herein. As used herein, a host 206 can be or can include various combinations of hardware and / or software, including a standalone server, a blade server, a cloud-implemented server, a distributed server, a virtual machine, a container, or a processing resource in a server farm. The host 206 can provide one or more services to one or more UEs.
[0090] The host 206 includes processing circuitry 502, which is operatively coupled to input / output interface 506, network interface 508, power source 510, and memory 512 via bus 504. Other components can be included in other embodiments. Features of these components can be substantially similar to those described for the devices of the previous figures, such as Figure 3 and Figure 4 ) so that their description generally applies to the corresponding components of the host 206.
[0091] Memory 512 can include one or more computer programs, including one or more host applications 514 and data 516, which can include user data, e.g., data generated by an encoder or decoder for the host 206 or data generated by the host 206 for a UE. Embodiments of the host 206 can utilize only a subset of the components shown or all of them. The host applications 514 can be implemented in a container-based architecture and can provide support for video codecs (e.g., Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), Motion Picture Experts Group (MPEG), VP9) and audio codecs (e.g., FLAC, Advanced Audio Coding (AAC), MPEG, G.711, EVS, Immersive Voice and Audio Services (IVAS)), including transcoding to multiple different categories, types, or implementations of UEs (e.g., cellphones, desktop computers, wearable display systems, heads-up display systems). The host applications 514 can also provide user authentication and permission checks and can periodically report health, routing, and content availability to a central node, such as a device in the core network or at the edge. Thus, the host 206 can select and / or indicate different hosts for over-the-top services for UEs. The host applications 514 can support various protocols, such as the HTTP Live Streaming (HLS) protocol, the Real-Time Messaging Protocol (RTMP), the Real Time Streaming Protocol (RTSP), Dynamic Adaptive Streaming over HTTP (MPEG-DASH), and the like.
[0092] Figure 6is a block diagram illustrating a virtualization environment 600 in which functions implemented by some embodiments of the audio encoder 202 or components of the audio encoder 202, or functions implemented by some embodiments of the audio decoder 212 or components of the audio decoder 212, can be virtualized. In current context, virtualization means creating virtual versions of devices or apparatuses which can include virtualizing hardware platforms, storage devices and network resources. As used herein, virtualization can apply to any device or component thereof described herein and involves an implementation in which at least a portion of the functionality is implemented as a virtual component. Some or all of the functionality described herein can be implemented as virtual components executed by one or more Virtual Machines (VMs) implemented in one or more virtualization environments 500 hosted by one or more hardware nodes such as a hardware computing device operating as a decoder, an encoder, a network node, a UE, a core network node, or a host. Further, in embodiments where a virtual node does not require radio connectivity (for example, a core network node or a host), then the node can be entirely virtualized.
[0093] The application 602, which can alternatively be referred to as a software instance, virtual appliance, network function, virtual node, virtual network function, etc., is run in the virtualization environment 600 to implement some of the features, functions, and / or benefits of some of the embodiments disclosed herein.
[0094] The hardware 604 comprises processing circuitry, memory storing software and / or instructions executable by the hardware processing circuitry, and / or other hardware devices as described herein, such as network interfaces, input / output interfaces, etc. The software can be executable by the processing circuitry to instantiate one or more virtualization layers 600 (also referred to as hypervisors or Virtual Machine Monitors (VMMs)), provide VMs 608A and 608B (one or more of which can be referred to generally as a VM 608), and / or implement any of the functions, features and / or benefits of some of the embodiments described herein. The virtualization layer 600 can present a virtual operating platform that appears like networking hardware to the VMs 608.
[0095] The VMs 608 comprise virtual processing, virtual memory, virtual network or interfaces and virtual storage, and can be run by a corresponding virtualization layer 600. Different embodiments of the instance of the virtual appliance 602 can be implemented on one or more of the VMs 608, and can be implemented in different ways. Virtualization of the hardware is sometimes referred to as Network Function Virtualization (NFV). NFV can be used to consolidate many network devices onto industry standard high-volume server hardware, physical switches and physical storage, which can be located in data centers, and customer premise equipment.
[0096] In the context of NFV, VMs 608 can be software implementations of physical machines that run programs as they would execute on a physical, non-virtualized machine. Each VM 608, along with the portion of hardware 604 that executes that VM, whether hardware dedicated to that VM and / or hardware shared by that VM with others, forms an individual virtual network element. Still in the context of NFV, a virtual network function is responsible for handling a particular network function that runs in one or more VMs 608 on top of hardware 604, which corresponds to application 602.
[0097] Hardware 604 can be implemented within a standalone network node with general or specific components. Hardware 604 can implement some functions via virtualization. Alternatively, hardware 604 can be part of a larger cluster of hardware, e.g., like in a data center or CPE, where many hardware nodes work together and are managed via management and orchestration 610 which is responsible for overseeing life cycle management of applications 602. In some embodiments, hardware 604 is coupled to one or more radio units, each of which includes one or more transmitters and one or more receivers that can be coupled to one or more antennas. Radio units can communicate with other hardware nodes directly via one or more appropriate networks, and can be used in combination with virtual components to provide a virtual node with radio capabilities, such as a radio access node or a base station. In some embodiments, some signaling can be provided using control system 612, which can optionally be used for communication between hardware nodes and radio units.
[0098] Turning now to Figure 7 A stage 1 (e.g., stage #1) FD-CNG VQ search using an offline created codebook 104 (e.g., as indicated above and shown in Figure 1
[0099] In some embodiments, in block 700, encoder 202 (i.e., as shown in Figure 2 For example, an input target vector of length N_target can be obtained using a 3GPP EVS codec analysis algorithm, as described in the reference 3GPP TS 26.445 encoder.
[0100] In some instances, the FD-CNG noise estimator relies on a hybrid spectral analysis approach. The low frequencies corresponding to the core bandwidth are covered by a high-resolution fast Fourier transform (FFT) analysis, while the remaining higher frequencies are captured by a CLD FB that exhibits a significantly lower spectral resolution of 400 Hz.
[0101] The input signal to the EVS audio / speech encoder can be configured to quantize the spectrum in the FD_CNG domain such that the EVS algorithm can quantize FDCNG target vectors of dimensions 17, 20, 21, and 24. The present disclosure describes the case when N_target is 24, however the first stage MSVQ method can also be applied to other dimensions. In alternative embodiments, N_target can be 21 (e.g., 21(N_WB)) without departing from the scope of the disclosed subject matter. However, special care should be taken when using the same stored codebook (e.g., a set of DCT-II trained codebooks trained on dimension 24) for other smaller dimensions (e.g., 17, 20, 21) to generate nearly distortionless target vectors of dimension 24 (e.g., same as the DCT dimension used in codebook training).
[0102] Exemplary operations can include a first operation comprising: representing a non-normalized target signal in the EVS description as N FD-CNG (i) (where 'N' stands for noise). Furthermore, in the EVS specification, N FD-CNG (i) is encoded in "5.6.3.5 Encoding of SID frames in FD-CNG". Notably, the length L SID corresponds to the variable N_target disclosed herein. N FD-CNG The final determination of the variable can be found in the "5.6.3.3 Adjustment of the first SID frame in FD-CNG" section (Eq. 1395) which targets the first SID frame after a speech frame.
[0103] The second operation includes both a log-domain conversion and a spectral envelope normalization. In the EVS specification, in "Section 5.6.3.5" equation (1396), N FD-CNG (i) signal is converted to the log10 dB domain. Then, the normalized signal In this description, the target signal target[i] of length N_target corresponds to the EVS specification signal SID of length L
[0104] In some embodiments, in operation 702, the encoder 202 removes a global offset value vector, denoted herein as the median vector. In other embodiments, the offset value vector can be a mean vector, a median vector, etc. This is done to reduce the dynamics of the signal in quantization and storage. In the FD-CNG domain, the obtained target signal is subtracted by the global offset value vector analyzed offline. An example of this operation can be mathematically described as:
[0105] target mr[i] = target[i] - midQ[i], for i e 0... (N_target - 1).
[0106] In some embodiments, in operation 704, the encoder can utilize DCT and IDCT related operations to add a global scaling factor to amplify the target vector to a dynamic and / or range maximized search domain. More specifically, to maximize storage and search precision within the Word8 mantissa, a global scaling factor is applied to the intermediate removed target vector. An example of this operation can be mathematically described as:
[0107] target mr_scaled[i] = target mr[i] * dct_invScaleF[1], for i e 0... (N_target - 1)
[0108] where dct_invScaleF[1] maximizes the storage dynamic and further amplifies the target vector to a search domain that includes 4 binomials or bits (i.e., 4 "decimals" in binary domain). The purpose of this DCT implementation related operation is to amplify the intermediate removed signal to a target level prior to applying the DCT to improve search precision (e.g., granularity). Without this amplification, the dct_target vector (after DCT processing) would exhibit a lower dynamic range. In floating point implementations, the lack of amplification is largely inconsequential, but in fixed point implementations (i.e., with finite precision), it is important to maximize the DCT input signal dynamic range and DCT output signal range. Otherwise, the DCT transform can add analysis noise without amplification.
[0109] In some embodiments, in operation 706, the scaled target vector signal is transformed by the encoder 202 to the DCT 24 search domain. The transformation results in a DCT target vector, denoted herein as "dct_target". The DCT Type II transform (as described in Figure 7 operation 706 in the DCT 24 search domain can be applied as follows:
[0110] dct_target = dct(target mr_scaled)
[0111] where target mr_scaled always has dimension 24 (NMAX_FDCNG). Even if N_target is lower than NMAX_FDCNG, the analysis DCT is performed using dimension NMAX_FDCNG. This avoids implementing multiple different DCTs (and IDCTs), and the mantissa and scaling factor can be optimized for one DCT length.
[0112] In some embodiments, the encoder 202 computes the common MSE contribution of globally truncated DCT coefficients. For DCT truncation lengths in all segments in the first stage codebook, potential common truncations in the DCT search domain yield a common high coefficient error contribution. If one additional first stage segment does not employ a transform to provide an adequate comparison of first stage mean squared error, then the error needs to be computed. An example of such an operation can include optionally synthesizing the truncated target and computing the common error, i.e., setting the DCT target (e.g., dct_target) vector components from max(Nseg) == 18 and / or truncating to N_target - 1 == 23 to zero.
[0113] In operation 112, the encoder 202 implements a pair-wise inner loop search to establish 8 (i.e., 4 x 2) "winners" from 4 pair-wise searches. The pair-wise search is a globally suboptimal pair-wise search over the four different codebooks of the codebook 104, followed by a post-optimization in operation 116.
[0114] To implement the search, the encoder 202 can establish a segment loop setup. The operations and / or steps described below are performed in this particular order for segm = 0 up to and including segm = Ns - 1. For example, in a first example step, tLen is set to truncLen[segm] for the duration of segment segm. Further, stl_mse_bear[segm][0] is set to a very large value (MAX FLOAT), and stl_mse_pair[segm][l] is set to a very large value (MAX FLOAT). This initialization ensures that these two values will be updated. Further, p_max, which points to the worst vector in the segment pair, is initialized to 0.
[0115] As part of the disclosed operations / steps, the encoder can compute the common MSE contribution of the truncated coefficients of a segment. The truncated error of each segment up to truncLen[Ns - 1] is needed in order to be able to finally compare the MSE error between different segments. For example, a first operation can include initializing an MSE variable by setting mse_trunc_segm[segm] = mse_trunc_all_segms. Further, the error energy of the truncated target coefficients of this segment can be summed over the global maximum truncation length truncLen[Ns - 1] >= tLen. Further, for i e 0... (truncLen[Ns - 1] - tLen - 1), mse_trunc_segm[segm] += (dct_target[tLen + i]) 2 .
[0116] As part of the disclosed operations / steps, the encoder 202 points to the codebook and scaling factors for the current segment. An example first operation and / or step includes: pointing to the current segment cb (the common codebook for the segment), where all vectors have a truncation length of truncLen[segm], where cb = cb_segmW8[segm]. In a subsequent step / operation, pointing to the current segment col_shift vector, where the coefficient column scaling codebook for the segment segm, where all col_shift vectors have a length of tLen, which can be mathematically represented as dct_col_shift_tab = col_shift[segm]. As used herein, the term "column shift" or "col_shift" can refer to an integer indicating an exponential value of a scaling factor, for example, equal to 2 [col_shift] The scaling factor of .
[0117] As part of the disclosed operations / steps, the encoder 202 implements a per-vector setup within each segment segm. For idx=0 and up to and including idx=nSeg[segm-1], the steps of summing the MSEs of the non-truncated coefficients are performed in this particular order. For example, a first example operation may include computing idx_full so that the MSE of the CB vector can be stored in a structured manner for each of the Nc candidates in the post-optimization step. Notably, idx_full=idx+nSegCum[segm]. The local mse is then initialized with the common contribution of the segment, i.e., mse idx = mse_trunc_segm[segm]. Then, for c∈0…(tLen-1), use tmp[c]=dct_target[c]–cb[idx][c]*2 col_shift_tab[c] Calculate the MSE for the current codebook vector idx, and for c∈0…(tLen-1), set mse idx +=tmp[c] 2 It is worth noting that it is the inner MSE calculation loop where the DCT truncation pays off in terms of reducing the WMOPS complexity, since now only coefficients in the range 0 to tLen-1 contribute to the MSE summation.
[0118] In some embodiments, the encoder 202 then saves the MSE value for the current vector index. In some embodiments, this step is optional and used as an "extended candidate" analysis or post-analysis step for selecting the final candidate vector from stage 1. For example, st1_mses[idx_full] = mse idxto set the MSE value. As indicated previously, these stored MSE values can be used for subsequent low complexity post-analysis steps.
[0119] In some embodiments, the encoder 202 then conditionally updates the best value pair for the segment. For example, a first example operation and / or step includes first evaluating whether the current vector is better than the worst value in the segment pair (i.e., whether the MSE of the current vector is lower than the highest MSE in the segment vector pair that has been evaluated so far). As used herein, a vector is "better" if it has a lower MSE, and a vector is "worst" if it has the highest MSE. Notably, the conditional update can start from the worst index that p_max points to,
[0120] if (mse idx <st1_mse_pair[segm][p_max]){
[0121] st1_idx_pair[segm][p_max] = idx_full;
[0122] }
[0123] In a second operation, the best MSE can be conditionally updated by:
[0124]
[0125] In a third operation, p_max is updated after each new candidate. It is always re-evaluated whether the new worst candidate is stored.
[0126]
[0127] Here, the old MSVQ stage #1 solution keeps a complex bookkeeping record, e.g., a list of 8 (or even 24) values, which potentially leads to a large worst case (WC) complexity problem.
[0128] At this stage, the encoder 202 has processed all Ns segments.
[0129] Notably, there are Ns*2 stored preliminary candidates in the variables:
[0130] st1_mse_pear[Ns][2] (best mse values from the segment-wise search)
[0131] st1_idx_part[Ns][2] (best idx_full values from the segment-wise search)
[0132] where Nv(128) stored MSE values in stl_mse have been created.
[0133] In this example with Ns=4, the required Nc=8 stage#2 candidate values have been determined. Out of these Nc values, at least two pairs of values from the serial pair-wise search are known to correspond to the global minimum out of the 8 values. However, Nc-2 (i.e. 6 in this case) values can represent sub-optimal candidates for stage#2 of the multi-stage VQ.
[0134] As the DCT type II truncated segments have different truncation lengths, these segments will differ in high frequency component content. In other words, it can be considered that different codebook segments apply different low-pass filters.
[0135] Typically, these Ns*2=Nc candidates can be accepted for the second stage search. However, in the case where the input target signal has more than one good vector within a segment, it is beneficial to check whether some of the candidates (e.g. the final pair-wise search candidate vectors) can be replaced by better candidates.
[0136] As a preparation for the second stage, the stl_mse_pair matrix is serialized into an MSE vector dist of length Ns*2, and the candidate index vector stl_idx_pair matrix is also serialized in the same way into a vector indices of length Ns*2.
[0137] (Note: this in-place serialization is a no-cost operation in C)
[0138] In some embodiments, a pointer p_max is updated to point to the worst candidate in dist and indices by searching for the maximum MSE value in dist.
[0139] For example, the first operation can comprise updating a pointer p_max to point to the worst candidate in the vector dist and indices by searching for the maximum value in dist p_max = maximum(dist, Nc), where max_index = maximum(vec, len) is a function that selects the maximum index for the values in a vector vec of length len and returns the index of the maximum value in vec as max_index.
[0140] Furthermore, the positions of the two best candidates (i.e. lowest MSE) are as follows:
[0141] p_min[0] = minimum(dist, Nc);
[0142] mse_memory = dist[p_min[0]]; / * remember best MSE value * /
[0143] dist[p_min[0]] = MAX_FLOAT; / * very high value * /
[0144] p_min[1] = minimum(dist, Nc);
[0145] dist[p_min[0]]) = mse_memory / * restore * /
[0146] where min_index = minimum(vec, len) is a function that selects the smallest index of the values in a vector vec of length len and returns the index of the smallest value in vec as min_index.
[0147] In operation 116, the encoder 202 uses the cyclic MSE neighbor index list to select the final set of Nc candidates.
[0148] In some embodiments, the cyclic MSE neighbor index list is implemented by the encoder 202 in operation 708. Offline, the vectors in the global temporary codebook (e.g., the serial concatenation of the vectors in the codebook cb_temporary_full = [cb_segmW8[0], cb_segmW8[1], cb_segmW8[2], cb_segmW8[Ns-1]] in this case, where Ns = 4) have been analyzed (e.g., offline analysis) as a self-MSE ordered vector of indices of length Nv.
[0149] In some embodiments, this cyclic ordered vector version can be obtained by analyzing the MSE between the codebook vectors and applying an approximate traveling salesman (TSP closed) problem solution to the vectors in cb_temporary_full. The distance measurement between two cities corresponds to the MSE between two individual vectors in the codebook. Essentially, each vector index in the codebook represents a city for a traveling salesman problem statement, and the MSE between the vectors corresponds to the distance between two cities.
[0150] In this example, a convex hull approach for the closed TSP problem is used to obtain an ordered vector of length Nv across all segments in the full concatenated codebook cb_temporary_full, mse_order_circ. To save on the cyclic complexity at the cost of finite ROM cost, the vector of cyclic ordered indices mse_order_circ can be used to create two auxiliary vectors neighb_mse_fwd and neighb_mse_rev, as shown in Table 3 below.
[0151]
[0152] Table 3
[0153] Notably, neighb_mse_fwd and neighb_mse_rev can also be created during runtime using the indexed mse_order_circ vector.
[0154] As part of the optimization in operation 116 (e.g., in Figure 1 or Figure 7 ), the encoder 202 can implement first operations and / or steps, including setting a vector check_ind of a total of Ncheck = 8 promising cross-segment candidate indices (full indices):
[0155] / * Use the two best candidates so far from the MSE circular neighbors * /
[0156] check_ind[0] = neighb_mse_fwd[p_min[0]];
[0157] check_ind[1] = neighb_mse_rev[p_min[0]];
[0158] check_ind[2] = neighb_mse_fwd[p_min[1]];
[0159] check_ind[3] = neighb_mse_rev[p_min[1]];
[0160] Afterwards, second operations include implementing one additional step in each direction of the circular list:
[0161] check_ind[4] = neighb_mse_fwd[check_ind[0]];
[0162] check_ind[5] = neighb_mse_rev[check_ind[1]];
[0163] check_ind[6] = neighb_mse_fwd[check_ind[2]];
[0164] check_ind[7] = neighb_mse_rev[check_ind[3]];
[0165] In a third operation, the saved global MSE values are used to check whether these are better candidates than the initial Nc candidates from the fast pair-wise search. In some instances, at most Nc-2 candidates are replaced (e.g., in steps described below).
[0166] For example, the two best candidates are first excluded (note: global exclusion is used for not reselecting / using the two (best) MSE values so far).
[0167] st1_mses[p_min[0]] = FLT_MAX; / * exclude * /
[0168] st1_mses[p_min[1]] = FLT_MAX; / * exclude * /
[0169]
[0170] In a fourth operation, there can be a first stage search part for the second stage candidate index, while Nc candidates are available in indices[], and their MSEs are available in dist of length Nc.
[0171] As shown in FIG. 11, the results of operation 116 are provided to operation 118 by encoder 202. The final candidates are transformed to the original FD-CNG domain by encoder 202. Figure 7 In some embodiments, the second stage requires to have the correct error signal in the FDCNG domain (e.g., each of the Nc vectors can have a corresponding error vector signal res[cand] also known as residual signal as input to the 2nd stage), as the upper stage has already scaled and prepared in that particular domain.
[0172] For example, the higher FD-CNG stage can be found in EVS 3GPP specification TS 26.445.
[0173] It is noted that since the DCT transform / rotation does not change the MSE distance to the target signal, it is not necessary to recompute the best Nc MSE values for the candidates. However, if the input signal length is shorter than the search (and stored codebook) DCT length, the Nc best MSEs can be updated to the appropriate shorter MSE domain by excluding the errors from the upper zeroed or upper extended part. In operation 710, the following steps are implemented by encoder 202. It is noted that the first operation and / or steps include, for each of the selected Nc candidates in indices[], obtaining the residual vector signal res[cand][
[0174] It is noted that since the DCT transform / rotation does not change the MSE distance to the target signal, it is not necessary to recompute the best Nc MSE values for the candidates. However, if the input signal length is shorter than the search (and stored codebook) DCT length, the Nc best MSEs can be updated to the appropriate shorter MSE domain by excluding the errors from the upper zeroed or upper extended part. In operation 710, the following steps are implemented by encoder 202. It is noted that the first operation and / or steps include, for each of the selected Nc candidates in indices[], obtaining the residual vector signal res[cand][
[0175]
[0176] In a second operation, a stage#2 (i.e. second stage) search can be started, where Nc candidate indices are available in indices[], their MSEs are available in dist[], and the residual signal in the FDCNG domain is available as res[Nc][N_target], and the worst candidate is denoted by p_max.
[0177] The encoder 202 (or decoder 212) receives an index idx_full in the range [0…(Nv-1)], and reconstructs the transmitted FDNCG vector fdcng_final of length N_target as follows idx_fall .
[0178] In a first operation, idx_full is decomposed into a segment segm and a local idx, i.e. first the segment segm is identified, and the local index idx.
[0179] segm = 0;
[0180]
[0181] idx = idx_full - nSegCum[segm];
[0182] idx is now the local index in the segment-specific codebook cb_segmW8[segm].
[0183] In a second operation, the algorithm points to the correct exponent table for segment segm by:
[0184] expTable = col_shift[segm];
[0185] Note: the length of the expTable vector is tLen = truncLen[segm].
[0186] In a third operation, the process includes retrieving and scaling the FDCNG vector with the DCT domain coefficients corresponding to idx from the table ROM memory
[0187]
[0188] It is noted that the mantissa values cb_segmW8[segm][idx][c] stored as bytes (Word8) can be retrieved and scaled using the c binary left shift operator “<<”. In full floating point operations, the dct_vec vector is equivalently obtained as:
[0189] dct_vec[c] = (cb_segmW8[segm][idx][c])*2 expTable[c] , for c e 0... (tLen - 1)
[0190] Optionally, in the optional DSP-optimized fixed-point implementation, the DSP instruction set can not be able to efficiently retrieve Word8 from ROM or RAM. To manage this case, the coefficients in consecutive byte pairs (Wordl6) can preferably be retrieved as follows:
[0191] Retrieve and scale the FDCNG vector with DCT-domain coefficients corresponding to idx from the table ROM memory:
[0192]
[0193] where AND(a,b) is the bitwise "and" function operating on 16-bit (Wordl6) variables. For example, the mantissa value cb_segmW8[segm][idx] is retrieved as a 2-byte block (Wordl6) in this DSP-optimized version. In some embodiments, each word is masked with a bitwise AND and appropriately scaled using the C binary left shift operator " << ". This DSP optimization can be applied to both the first-stage search in the encoder and the first-stage reconstruction in the decoder.
[0194] In a fourth operation, the DCT-domain vector is transformed to the intermediate removed FDCNG domain by setting the dct_vec vector components above tLen to zero, such that
[0195] dct_vec[i] = 0.0; for i e tLen... (NMAX_FDCNG - 1).
[0196] Further, the IDCT inverse DCT Type II transform is applied as follows:
[0197] idct_vecl = idct(dct_vec); where dct_vec has dimension 24 (NMAX_FDCNG).
[0198] Notably, even if N_target is lower than NMAX_FDCNG, the synthesis IDCT runs with dimension NMAX_FDCNG.
[0199] In a fifth operation, the vector is scaled down to the original unscaled FDCNG domain (e.g., using a global value to up-scale the Word8 mantissa representation) by:
[0200] idct_vec_global_scaled[i]=idct_vec1[i]*dct_scaleF[1],for,for i∈0…(N_target-1).
[0201] In the sixth operation, the intermediate stage#1CB vector is added by:
[0202] fdcng_final idx_full [i]=idct_vec_global_scaled[i]+midQ[i], for i∈0…(N_target-1).
[0203] In some embodiments, this reconstruction principle enables fast search in the residual (offset removed) truncated DCT-II domain and further enables storage of CB code vectors of mantissa size Word8.
[0204] It is worth noting that the codebook 104 created for dimension 24 (NMAX_FDCNG) can be used for vectors of different dimensions. Figure 8 and Figure 9 , where the codebook 104 is used to quantize a vector of dimension 21 (N_WB). Figure 8 is a block diagram of a system view of FD-CNG-VQ with low ROM and low WMOPS supporting shorter input target vectors. Figure 8 Similar to Figure 1 , but with two additional blocks, namely block 800 and block 802. Unless otherwise specified, the encoder 202 is Figure 8 The blocks 102 to 122 are implemented as described above. Figure 1 The same operations are performed as in the corresponding blocks in [ ]. In block 801, the FDCNG 21 domain values of the target vector are extrapolated to FDCNG 24 domain values. As previously described, the output of block 120 has values corresponding to the dimensional DCT24 domain. In block 802, the DCT24 domain MSE values are updated to the FDCNG 21 domain.
[0205] Now refer to Figure 9 We will now describe the first stage (ie, stage #1) FD-CNG VQ search using codebook 104 that was created offline for dimension 24 (NMAX_FDCNG) but is used to quantize vectors of dimension 21 (N_WB). Figure 9 Similar to Figure 7 , but with two additional blocks, namely block 900 and block 902. Unless otherwise specified, the encoder 202 is Figure 9 The blocks 104, 110-120, 122 and 700-704 shown are implemented in the same manner as described above. Figure 7the same as the corresponding block in
[0206] At block 700, the encoder 202 obtains a target vector of dimension 21 (e.g., a shorter dimension than used for codebook training). The input target vector of length N_WB can be obtained using the 3GPP EVS codec analysis algorithm as described in reference 3GPP TS 26.445 encoder.
[0207] In some embodiments, in operation 900, the input target_wb signal of dimension N_WB is converted by the encoder 202 to the FDNCG 24 input domain by a fine extrapolation method to dimension N_MAX_FDCNG. The extrapolation results in an extended input target of length N_MAX_FDCNG.
[0208] In a first operation, the FDNCG target_wb signal is extrapolated in the FDNCG domain by applying a shorter N = 21 DCT Type II transform as follows:
[0209] dct_target_wb = dct(target_wb);
[0210] where target_wb has dimension 21 (N_WB), and the applied DCT function is for the same dimension, i.e., since N_WB is lower than NMAX_FDCNG, an additional analysis DCT is implemented with dimension N_WB (= 21). This is to avoid storing multiple different DCT domain FD-CNG codebooks. Thus, the stored mantissa and scaling factor values can be optimized for one single larger DCT length (e.g., 24), and ROM space is saved.
[0211] In this example, only the DCT-II (N_WB) is performed up to obtain coefficients 0 to (Ntr_WB - 1), where Ntr_WB is 18. This is to reduce the complexity of the additional DCT-II (N = 21) analysis and subsequent extrapolation. A truncated length other than Ntr_WB = 18 can be used in this DCT analysis operation by not analyzing the high frequency components (e.g., the upper basis vector). However, it is relevant to provide slightly more detailed information in the target extrapolation than the truncation used for the stored DCT (24) codebook.
[0212] In this example, the DCT-II (N_WB = 21) analysis is truncated after 18 coefficients (i.e., 85.7% of the full bandwidth), while the stored DCT- II (N = 24) vector is truncated at 75% of the bandwidth.
[0213] Reference Figure 10DCT-II (N=21) is shown as a set of DCT coefficients (i.e. scaling factors for the cosine basis vectors used for the DCT) with Ntr_WB coefficients in dct_target_wb as input signal target_wb. For example, dct_target_wb[0] is the scaling of the DC basis vector 'BAS[0]', dct_target_wb[6] is the scaling of the third basis vector 'BAS[6]' as shown in Fig. 3. There are also DCT-II (N=21) basis vectors (which are used for the extrapolation) available, either as stored in the transform matrix bas(k,t), or alternatively by dynamically constructing the basis vectors using the IDCT (type II, N=N_WB) formula as follows: Figure 10
[0214]
[0215] Each available basis vector can be accessed as follows:
[0216] bas(k), k e 0... (Ntr_WB - 1), specifies the k-th basis vector, where bas(k,t) indexes the entire matrix using the time index element t e 0... (N_WB - 1) for each bas(k) basis vector.
[0217] Furthermore, now by using a subset of the DCT basis vectors for the extrapolation, the input signal target_wb is extended by (N_MAX_FDCNG - N_WB) samples and each basis vector bas(k) is scaled by the corresponding DCT coefficient dct_target_wb(i) and then the extended part of the basis vectors is summed up.
[0218] This extension / extrapolation ensures that the extended target_wb signal target_ext will preserve the dominant frequency components in target_wb without adding unwanted and unnecessary extension / extrapolation noise which would lead to a degradation of the overall VQ performance.
[0219] Other extrapolation methods like zero extension or last value repetition, interpolation or linear extrapolation or polynomial extrapolation do not preserve the cosine waveform content to the same extent after the DCT-II (24) analysis. It is important to preserve the undistorted tones in the extended signal as an extended cosine, because the subsequent transform step (DCT-II (N=24)) is also based on a cosine analysis.
[0220] Furthermore, in order to save complexity, the target_wb signal can be reused for the initial part of the extended signal target_ext as follows:
[0221] target_ext(t) = target_wb(t), t e 0... (N_WB - 1)
[0222] Optionally, the initial portion can be recreated with IDCT-II (N = N_WB) by using up to coefficient (Ntr_WB - 1) of the available DCT coefficients.
[0223] Furthermore, the creation of the extension by scaling the basis vector sum can be implemented. Due to the reflection property of the DCT II type basis vectors, only the higher part of the basis vectors is actually used for the proposed extrapolation. In some embodiments, this can be implemented by the following:
[0224]
[0225]
[0226] In the final operation, the extension target signal is copied to the full length target buffer as follows:
[0227] target(t) = target_ext(t), t e 0... (N_MAX_FDCNG - 1).
[0228] In some embodiments, the encoder 202 can now use the extended input signal target(t) to continue the search of the DCT-II (N = 24) domain codebook in the manner described earlier in block 702.
[0229] In Figure 10 , a method for extrapolation of the first seven unscaled DCT-II (N = 21) basis vectors is outlined.
[0230] It should be noted that even though the extrapolation described here is implemented in the incoming input normalized FDCNG envelope signal domain, the extrapolation can be equivalently implemented after the intermediate subtraction and global scaling (i.e., after operation 702 or after operation 704 in Figure 9 ). Notably, Figure 10 shows the initial extension and / or extrapolation of a number of basis vectors, followed by scaling using DCT 21 coefficients. These scaled basis vectors are then summed to generate the extended portion of the target as output.
[0231] The search process continues as previously Figure 7 described in Figure 9 , for the N_MAX_FDCNG sized target vector, then proceeds fully to block 902 in , where the full vector length of N_MAX_FDCNG can be used to implement the IDCT synthesis in order to be able to compute the upper extension portion error contribution.
[0232] In Figure 9 In block 710, residual calculation in preparation for stage#2 can be implemented on the shorter N_WB number of samples.
[0233] In block 902, the MSE for each candidate is updated to correctly reflect the reduced dimension N_WB mean squared error (i.e., core DCT-II (N=24) search, where the MSE reflects the error in the extended DCT-II domain (N=24), and needs to be updated for stage#2 and the forward search). If there is a new worst candidate after the MSE update for dimension N_WB, the p_max index of the worst candidate is also updated.
[0234] In some embodiments, the MSE stored in the dist[Nc] vector is updated by subtracting the MSE contribution ext_err[c] for the extended portion, as follows:
[0235]
[0236]
[0237] Further, p_max can be updated using p_max = maximum(dist, Nc). Notably, all 8 (Nc) of the new current worst candidates in the vector dist are established using the shorter dimension N_WB of MSE.
[0238] In some embodiments, the first stage (e.g., stage#1) search portion of the codebook preparation for second stage (e.g., stage#2) candidate indices stored using DCT-II (N=24) can be exited. Thus, Nc candidate indices are available in indices[], and their MSE for the shorter dimension N_WB are available in dist[Nc]. Further, p_max points to the worst candidate from the first stage.
[0239] Figure 11 An overall synthesis model usable for both the encoder and / or the decoder is shown. For example, Figure 11 The manner in which the encoder 202 reconstructs the FD-CNG vector in operation 120 is shown schematically. Turning to Figure 11 , the encoder 202 obtains an index idx_full within the range [0…(Nv-1)] 1100, which contains Ns-1 segments and corresponding column shift (e.g., described herein as ‘col_shift’) values 1102. In the case of the decoder, there would only be the fdcng_final idx_full .
[0240] In operation 1104, the encoder 202 or the decoder 212 implements the up-shift per segment and coefficient. In operation 1106, the encoder 202 or the decoder 212 implements the inverse DCT type II transform and downsizes the output to the original unscaled FDCNG domain in operation 1108.
[0241] In operation 1110, the intermediate stage #1 codebook vector (i.e., shown as the midQ
[24] vector) is added back to the vector to output and / or produce the fdcng_final idx_full vector.
[0242] According to some embodiments of the inventive concepts, the operations of the encoder 202 (e.g., implemented using the block diagram structure of Figure 12 will now be discussed with reference to the flowchart of Figure 3 For example, the modules can be stored in the memory 310 of the Figure 3 These modules can provide instructions to cause the encoder 202 to implement the respective operations of the flowchart when the instructions of the modules are executed by the corresponding encoder processing circuitry.
[0243] Figure 12 The operations implemented by the encoder 202 according to some embodiments are shown. Referring to Figure 12 In block 1201, the encoder 202 obtains a DCT target vector (e.g., the ‘dct_target’ vector). In some embodiments, the operations of block 1201 can be collectively represented by blocks 1301-1307 of Figure 13 More specifically, Figure 13 The operations implemented by the encoder 202 to obtain the dct_target vector according to some embodiments are shown. Referring to Figure 13 In block 1301, the encoder 202 obtains an input target vector. In block 1303, the encoder 202 removes a global offset value vector (e.g., the mid vector) from the input target vector to form an offset target vector. In block 1305, the encoder 202 applies a global scaling factor to the offset target vector. In block 1307, the encoder 202 transforms the offset target vector to a target discrete cosine transform search domain to form the dct_target vector.
[0244] As previously mentioned, the input target vector used to obtain the dct_target vector can have a different dimension than the dimension of the codebook 104. As Figure 14As shown in block 1401 of FIG. 1 , encoder 202 extrapolates the input target vector to the codebook dimensions. In some embodiments, encoder 202 extrapolates the input target vector to the codebook dimensions by extending the input domain basis vectors of the DCT transform used. In some embodiments, the extension of the input domain basis vectors of the DCT transform used is based on a subset of the input domain basis vectors.
[0245] Back to Figure 12 In block 1203, the encoder 202 performs a suboptimal pairwise inner search in each segment of a codebook 104 having a plurality of segments, each segment having a truncation vector different from truncation vectors of other segments in the plurality of segments, to determine a set of pairwise initial candidates from each segment in the plurality of segments, thereby forming a plurality of pairwise initial candidates from the suboptimal pairwise inner search.
[0246] Figure 15 An example embodiment of implementing a suboptimal pairwise inner search in codebook 104 is shown (e.g. Figure 12 1203). Go to Figure 15 , in block 1501 , for each segment in the plurality of segments, the encoder 202 initializes a segment pair to a value large enough to ensure that both values in the segment pair are updated.
[0247] In block 1503 , for each vector index in each segment, the encoder 202 determines whether the mean square error (MSE) of the vector index being analyzed is less than the MSE of the segment pair.
[0248] In block 1505, in response to the MSE of the vector index being less than the worst MSE of the segment pair, the encoder 202 updates the pairwise initial candidate set to include the vector index. In block 1507, in response to the MSE of the vector index being less than the best MSE of the segment pair, the encoder 202 updates the segment pair to include the vector index.
[0249] Optionally, in some embodiments, in response to the MSE of the current vector being higher than the worst MSE of the segment pair, the encoder 202 does not update the segment pair to include the vector index. Figure 12 In block 1205 , the encoder 202 (optionally) performs post-optimization on the plurality of pairwise initial candidates to replace the plurality of pairwise initial candidates in the plurality of pairwise initial candidates with the candidate neighbor lists, thereby forming a plurality of pairwise final candidates.
[0250] Figures 16A-1 6B shows an example embodiment of post-implementation optimization (e.g., Figure 12 1205). Go to Figure 16AIn block 1601, the encoder 202 determines which of the plurality of paired initial candidates has the lowest MSE of the plurality of paired initial candidates. For each paired initial candidate other than the paired initial candidate with the lowest MSE, the encoder 202 implements the operations of blocks 1603-1613.
[0251] In block 1603, the encoder 202 compares the MSE of the paired initial candidate to the MSE of the neighbor vector in the positive direction.
[0252] In block 1605, in response to the MSE of the paired initial candidate being higher than the MSE of the neighbor vector in the positive direction, the encoder 202 updates the paired initial candidate to the neighbor vector in the positive direction.
[0253] In block 1607, in response to the paired initial candidate updated to the neighbor vector in the positive direction having a lower MSE than the paired initial candidate with the lowest MSE, the encoder 202 sets the paired initial candidate updated to the neighbor vector in the positive direction as the paired candidate with the lowest MSE.
[0254] In block 1609, the encoder 202 compares the MSE of the paired initial candidate to the MSE of the neighbor vector in the negative direction. In block 1611, in response to the MSE of the paired initial candidate being higher than the MSE of the neighbor vector in the negative direction, the encoder 202 updates the paired initial candidate to the neighbor vector in the negative direction.
[0255] Turning to FIG. 16B, in block 1613, in response to the paired initial candidate updated to the neighbor vector in the negative direction having a lower MSE than the paired initial candidate with the lowest MSE, the encoder 202 sets the paired initial candidate updated to the neighbor vector in the negative direction as the paired candidate with the lowest MSE.
[0256] In some embodiments, the paired initial candidate is compared to the next neighbor vector in the positive direction and the negative direction. This is illustrated in Figure 17 Turning to Figure 17 In block 1701, the encoder 202 compares the MSE of the paired initial candidate to the MSE of the next neighbor vector in the positive direction.
[0257] In block 1703, in response to the MSE of the paired initial candidate being higher than the MSE of the next neighbor vector in the positive direction, the encoder 202 updates the paired initial candidate to the next neighbor vector in the positive direction.
[0258] In block 1705, the encoder 202 compares the MSE of the paired initial candidate to the MSE of the next neighbor vector in the negative direction.
[0259] In block 1707, in response to the MSE of the pair of initial candidates being higher than the MSE of the next neighbor vector in the reverse direction, the encoder 202 updates the pair of initial candidates to the next neighbor vector in the reverse direction.
[0260] In some embodiments, the neighbor vector and the next neighbor vector in the forward direction are part of a forward circular MSE neighbor index list, and the neighbor vector and the next neighbor vector in the reverse direction are part of a reverse circular MSE neighbor index list. The forward circular MSE neighbor index list and the reverse circular MSE neighbor index list are created based on the ordering of the vectors of length Nv across all the plurality of segments in the full temporary codebook cb_temporary_full, mse_order_circ.
[0261] Returning to Figure 12 In block 1207, the encoder 202 reconstructs the plurality of pair of final candidates using an inverse discrete cosine transform type II (DCT-II) to transform the plurality of pair of final candidates to the original domain.
[0262] As previously mentioned, the input target vector used to obtain the dct_target vector can have a different dimension than the dimension of the codebook 104. As Figure 14 shown in block 1403, when this happens, the encoder 202 updates the plurality of reconstructed candidates to the dimension of the input target vector.
[0263] Figure 18 An example embodiment of reconstructing the plurality of pair of final candidates is shown. Turning to Figure 18 In block 1801, the encoder 202 obtains an index idx_full within the range [0…(N_v-1)] 1100, which contains the plurality of segments within the range of 0 to Ns-1, and the corresponding column shift value, such as the col_shift value 1102.
[0264] In block 1803, the encoder 202 implements the segment and coefficient based upshifting. In block 1805, the encoder 202 implements an inverse discrete cosine transform (DCT) type II transform. In block 1807, the encoder 202 down-scales the output vector from the DCT type II transform to the original unscaled FD-CNG domain vector. In some embodiments, the scaling can include upscaling or downscaling.
[0265] In block 1809, the encoder 202 adds the global offset value vector (e.g., the median and / or the intermediate stage #1 codebook vector) back to the unscaled FD-CNG domain vector to form the fdcng_final idx_full vector.
[0266] A similar reconstruction can be performed in the decoder 212. This is shown in Figure 19 Turning to Figure 19 In block 1901, the decoder 212 receives an index, which corresponds to a segment, a segment codebook vector (e.g., a mantissa vector), an associated column shift value, and a global offset value vector.
[0267] In block 1903, the decoder 212 performs an upshift operation on the segment codebook vector using the column shift value to form an upshifted vector. In block 1905, the encoder 202 performs an inverse DCT type II transform on the upshifted vector to produce an output vector. In block 1907, the decoder 212 down-scales the output vector from the inverse DCT type II transform to the original unscaled FD-CNG domain vector.
[0268] In block 1909, the decoder 212 adds the global offset vector (e.g., the intermediate vector) back to the unscaled FD-CNG domain vector to form a final output vector (e.g., fdcng_final idx_full vector and / or the final FDCNG domain vector).
[0269] Summary of ROM savings in the above encoder embodiment.
[0270] The total ROM Word8 storage in the proposal is the sum of (128 170 272 1404) = 1974 bytes.
[0271] The forward and backward circular list vectors for enhanced search each add 128 Word8, i.e., +256 bytes.
[0272] The total ROM storage for the original LBG codebook is 3072 Word16 = 6144 bytes
[0273] The total ROM storage for the original LBG codebook 3072 single precision floating point numbers = 9216 bytes.
[0274] The table ROM reduction is approximately (6144 - (1974 + 256)) / 6144 => approximately 63%.
[0275] Summary of WMOPS savings on loops (operations) in the above encoder embodiment.
[0276] The number of inner loop coefficients to process has been reduced by 30% (from 3072 to 1974).
[0277] The inner best MSE update loop (runs 128 times) has been optimized (using pairing) to use only 7-8 loops.
[0278] Notably, in the worst case, the reference uses approximately ~35 cycles to maintain and update the list of best Nc=8 candidates. Thus, in the worst case (WC), for the update of the best candidates, a saving of (35-7)=28 operations (or cycles) can be expected.
[0279] In this proposal, the MSE computation for each vector increases by 1 op * tLen due to the shift up of the coefficient mantissa, and the saving of the MSE per index costs 2 ops, but this increase (1 op * tLen + 2 ops) is always lower than the WC saving of the best MSE update loop (28).
[0280] When used for target vectors of different dimensions (e.g. N_WB), a codebook 104 created for a dimension of e.g. 24 (NMAX_FDCNG) can achieve similar inner loop MSE savings.
[0281] Although the computing devices described herein (e.g., decoders, audio object renderers, encoders, hosts) can include the illustrated combinations of hardware components, other embodiments can include computing devices with different combinations of components. It will be appreciated that these computing devices can include any suitable combination of hardware and / or software necessary to implement the tasks, features, functions, and methods disclosed herein. Determining, computing, obtaining, or similar operations described herein can be implemented by processing circuitry, which can process information, and make determinations, by, for example, transforming the obtained information, comparing the obtained information or transformed information to information stored in the network node, and / or implementing one or more operations based on the obtained information or transformed information, and making a determination as a result of the processing. Furthermore, while components are described as single boxes located within a larger box, or nested within multiple boxes, in practice the computing devices can include multiple different physical components constituting a single illustrated component, and the functionality can be divided among separate components. For example, a communications interface can be configured to include any of the components described herein, and / or the functionality of components can be divided between processing circuitry and a communications interface. In another example, non-computationally intensive functionality of any such components can be implemented in software or firmware, and computationally intensive functionality can be implemented in hardware.
[0282] In certain embodiments, some or all of the functionality described herein can be provided by a processing circuitry that executes instructions stored in a memory, which in certain embodiments can be a computer program product in the form of a non-transitory computer readable storage medium. In alternative embodiments, some or all of the functionality can be provided by a processing circuitry without executing instructions stored on a separate or discrete device readable storage medium, e.g., in hardware. In any of these particular embodiments, whether the processing circuitry executes instructions stored on a non-transitory computer readable storage medium or not, the processing circuitry can be configured to implement the described functionality. The benefits provided by such functionality can extend to the entire computing device, and / or to end users and wireless networks generally. Although the computing devices described herein (e.g., UEs, encoders, decoders, network nodes, hosts) can include the combination of hardware components illustrated, other embodiments can include computing devices with different combinations of components. It is to be understood that these computing devices can include any suitable combination of hardware and / or software presently known or developed in the future necessary to implement the tasks, features, functions, and methods disclosed herein. Determining, calculating, obtaining, or similar operations described herein can be implemented by a processing circuitry that can process information, e.g., by converting the obtained information into other information, comparing the obtained information or converted information to information stored in the network node, and / or implementing one or more operations based on the obtained information or converted information, and as a result of said processing making a determination. Furthermore, although components are described as single boxes or nested within multiple boxes within a larger box, in practice, a computing device can include multiple distinct physical components constituting a single illustrated component, and the functionality can be divided between the separate components. For example, a communications interface can be configured to include any of the components described herein, and / or the functionality of the components can be divided between the processing circuitry and the communications interface. In another example, non-computationally intensive functionality of any of such components can be implemented in software or firmware, and computationally intensive functionality can be implemented in hardware.
[0283] In certain embodiments, some or all of the functionality described herein can be provided by a processing circuitry that executes instructions stored in a memory, which in certain embodiments can be a computer program product in the form of a non-transitory computer readable storage medium. In alternative embodiments, some or all of the functionality can be provided by a processing circuitry without executing instructions stored on a separate or discrete device readable storage medium, e.g., in hardware. In any of these particular embodiments, whether the processing circuitry executes instructions stored on a non-transitory computer readable storage medium or not, the processing circuitry can be configured to implement the described functionality. The benefits provided by such functionality can extend to the entire computing device, and / or to end users and wireless networks generally.
[0284] A number of example tables used in the detailed description of various embodiments are presented below. For example, Table 4 illustrates example global scaling and global subtraction tables. Notably, the second row of Table 4 illustrates global scaling constants, while the third row illustrates an example offset vector.
[0285]
[0286] Table 4
[0287] Segmented codebook for MSVQ stage #1
[0288] Total ROM Word8 storage is the sum of (128 170 272 1404) = 1974 bytes.
[0289] In some embodiments, examples of the above-described segmented codebooks are illustrated in Table 5. Notably, these segmented codebooks can include "optimized" segmented codebooks that include a mantissa limit of 8, and an exponent that is not significantly limited. For example, these segmented codebooks can be signal-to-noise ratio (SNR) limited or granularity limited.
[0290]
[0291]
[0292]
[0293]
[0294] Table 5
[0295] The following Table 6 illustrates example segmented scaling factors.
[0296]
[0297] Table 6
[0298] Cyclic ordering vector
[0299] The following Table 7 illustrates an example MSE ordered cyclic neighbor relation table (i.e., mse_order_circ[]) for the above-described codebook. Notably, Table 7 was created using a convex hull method to solve a TSP problem (e.g., a closed loop TSP problem). Table 7 also includes neighb_mse_fwd, which is the forward direction neighbor vector mentioned above, and is the preferred way to traverse the cyclic MSE neighbor list in the forward direction. In addition, Table 7 also includes neighb_mse_rev, which is the reverse direction vector mentioned above, and is the preferred way to traverse the cyclic MSE neighbor list in the reverse direction.
[0300]
[0301]
[0302]
[0303] Table 7
Claims
1. A method implemented by an encoder, the method comprising: obtaining (1201) a discrete cosine transform (DCT) target vector; implementing (1203) a suboptimal pair-wise inner search in each segment of a codebook (104) having a plurality of segments to determine a pair-wise initial candidate set from each of the plurality of segments, forming a plurality of pair-wise initial candidates from the suboptimal pair-wise inner search, wherein each segment has a clipping vector that is different from the clipping vectors of the other segments of the plurality of segments, wherein the DCT target vector is a target of the suboptimal pair-wise inner search in each of the plurality of segments; reconstructing (1207) a plurality of final candidates using an inverse type-II discrete cosine transform (DCT-II) to transform the plurality of final candidates into final candidate data in an original domain; and providing (1209) the final candidate data to a second stage of a multi-stage vector quantizer.
2. The method of claim 1, wherein, the plurality of segments includes four segments.
3. The method of any one of claims 1-2, wherein, obtaining the DCT target vector includes: obtaining (1301) an input target vector; removing (1303) a global offset value vector from the input target vector to form an offset target vector; applying (1305) a global scaling factor to the offset target vector to form a scaled offset target vector; transforming (1307) the scaled offset target vector to a target discrete cosine transform search domain to form the DCT target vector.
4. The method of claim 3, further comprising: in response to the input target vector having a dimension that is different from a dimension of the codebook: extrapolating (1401) the input target vector to the dimension of the codebook; and wherein transforming the plurality of final candidates into a plurality of reconstructed final candidates in an original domain includes updating (1403) the plurality of reconstructed candidates to the dimension of the input target vector. extrapolating the input target vector to the dimension of the codebook includes extrapolating the input target vector by an extension of an input domain basis vector of the used DCT transform.
5. The method of claim 4, wherein, the extension of the input domain basis vector of the used DCT transform is based on a subset of input domain basis vectors.
6. The method of claim 5, wherein, implementing the suboptimal pair-wise inner search includes:
7. The method of any one of claims 1-6, wherein, for each segment of the plurality of segments, initializing (1501) a segment pair to a value large enough to ensure that both values in the segment pair will be updated; for each vector index in each segment: determining (1503) whether a mean squared error (MSE) of the vector index being analyzed is less than a MSE of the segment pair; in response to the MSE of the vector index being less than a worst MSE of the segment pair, updating (1505) the pair-wise initial candidate set to include the vector index; in response to a MSE of a current vector being less than a best MSE of the segment pair, updating (1507) the segment pair to include the vector index. implementing post-optimization on the plurality of pair-wise initial candidates includes:
8. The method of any one of claims 1-7, wherein, determining (1601) which pair-wise initial candidate of the plurality of pair-wise initial candidates has a lowest mean squared error (MSE) of the plurality of pair-wise initial candidates; for each pair initial candidate other than the pair initial candidate with the lowest MSE: compare (1603) the MSE of the pair initial candidate with the MSE of a neighbor vector in the forward direction; update (1605) the pair initial candidate to the neighbor vector in the forward direction in response to the MSE of the pair initial candidate being higher than the MSE of the neighbor vector in the forward direction; set (1607) the pair initial candidate updated to the neighbor vector in the forward direction as the pair candidate with the lowest MSE in response to the pair initial candidate updated to the neighbor vector in the forward direction having a lower MSE than the pair initial candidate with the lowest MSE; compare (1609) the MSE of the pair initial candidate with the MSE of a neighbor vector in the reverse direction; update (1611) the pair initial candidate to the neighbor vector in the reverse direction in response to the MSE of the pair initial candidate being higher than the MSE of the neighbor vector in the reverse direction; set (1613) the pair initial candidate updated to the neighbor vector in the reverse direction as the pair candidate with the lowest MSE in response to the pair initial candidate updated to the neighbor vector in the reverse direction having a lower MSE than the pair initial candidate with the lowest MSE.
9. The method of claim 8, further comprising: compare (1701) the MSE of the pair initial candidate with the MSE of a next neighbor vector in the forward direction; update (1703) the pair initial candidate to the next neighbor vector in the forward direction in response to the MSE of the pair initial candidate being higher than the MSE of the next neighbor vector in the forward direction; compare (1705) the MSE of the pair initial candidate with the MSE of a next neighbor vector in the reverse direction; update (1707) the pair initial candidate to the next pair vector in the reverse direction in response to the MSE of the pair initial candidate being higher than the MSE of the next neighbor vector in the reverse direction.
10. The method of any one of claims 8-9, wherein, the neighbor vector and the next neighbor vector in the forward direction are part of a forward circular MSE neighbor index list, and the neighbor vector and the next neighbor vector in the reverse direction are part of a reverse circular MSE neighbor index list.
11. The method of claim 10, wherein, the forward circular MSE neighbor index list and the reverse circular MSE neighbor index list are created based on a vector ordering mse_order_circ across all the plurality of segments in a full codebook cb_temporary_full of length Nv.
12. The method of any one of claims 1-11, wherein, reconstructing the plurality of pair final candidates comprises: obtaining (1801) an index idx_full in the range [0…(Nv-1)] containing a plurality of segments in the range 0 to Ns-1 and corresponding column shift values; performing (1803) a segment-wise and coefficient-wise upshift of the plurality of segments using the column shift values for the indices to form upshifted segments; performing (1805) an inverse discrete cosine transform, DCT Type II transform on the upshifted segments; scaling (1807) an output vector from the inverse DCT Type II transform to an original, un-scaled FDCNG domain vector; and The global offset value vector is added (1809) back to the unscaled frequency domain comfort noise generation FD-CNG domain vector to form fdcng_final idx_full vector.
13. The method of claim 1, wherein, the truncated vectors include vector segments having different degrees of high frequency content compared to other segments.
14. The method of claim 1, wherein, the final candidate data includes one or more of: a vector index, a reconstructed candidate vector, and / or a transformed codebook entry.
15. The method of claim 1, further comprising: performing (1205) post-optimization on the plurality of paired initial candidates to replace ones of the plurality of paired initial candidates using a candidate neighbor list to form a plurality of final candidates.
16. A method of reconstructing a target vector in a decoder (212), the method comprising: receiving (1901) indices corresponding to segments, associated column shift values, segment codebook vectors, and a global offset value vector; performing (1903) an upshift operation on the target vector using the column shift values to form an upshifted vector; performing (1905) an inverse discrete cosine transform, DCT Type II transform on the upshifted vector to produce an output vector; scaling (1907) the output vector from the inverse DCT Type II transform to an original, un-scaled frequency domain comfort noise generation, FDCNG, domain vector; and The global offset value vector is added (1909) back to the unscaled FD-CNG domain vector to form fdcng_final idx_full Vector.
17. An encoder (202) comprising: processing circuitry (302); and memory (310) coupled with the processing circuitry, wherein the memory includes instructions that, when executed by the processing circuitry, cause the encoder (202) to perform operations according to any of claims 1-15.
18. An encoder (202) adapted to perform a method according to at least one of claims 1-15.
19. An encoder (202) adapted to: obtain a discrete cosine transform, DCT, target vector; implementing a suboptimal pair-wise inner search in each segment of a codebook (104) having a plurality of segments to determine a set of pair-wise initial candidates from each segment of the plurality of segments to form a plurality of pair-wise initial candidates from the suboptimal pair-wise inner search, wherein, each segment has a truncated vector of the truncated vectors that is different from the truncated vectors of other segments in the plurality of segments, wherein the DCT target vector is a target of the sub-optimal paired intra search in each segment of the plurality of segments; reconstruct a plurality of final candidates using an inverse II-type discrete cosine transform, DCT-II, to transform the plurality of final candidates to final candidate data in an original domain; and provide the final candidate data to a second stage of a multi-stage vector quantizer.
20. The encoder of claim 19, wherein, the plurality of segments includes four segments.
21. The encoder of claim 19 or 20 adapted to obtain the DCT target vector by: obtaining an input target vector; removing a global offset value vector from the input target vector to form an offset target vector; applying a global scaling factor to the offset target vector to form a scaled offset target vector; transforming the scaled offset target vector to a target discrete cosine transform search domain to form the DCT target vector.
22. The encoder of claim 21, further adapted to: in response to the input target vector having a different dimensionality than a dimensionality of the codebook: extrapolate the input target vector to the dimensionality of the codebook; and wherein, transforming the plurality of final candidates to a plurality of reconstructed final candidates in an original domain includes updating a plurality of reconstructed candidates to a dimensionality of the input target vector.
23. The encoder of claim 22, wherein, extrapolating the input target vector to the dimensionality of the codebook includes extrapolating the input target vector by an extension of input domain basis vectors of the DCT transform used.
24. The encoder of claim 23, wherein, the extension of input domain basis vectors for the DCT transform is based on a subset of input domain basis vectors.
25. The encoder of any of claims 19-24, wherein, implementing the suboptimal pair-wise inner search includes: for each segment of the plurality of segments, initializing a segment pair to a value large enough to ensure that both values in the segment pair will be updated; for each vector index in each segment: determining whether a mean squared error (MSE) of the vector index being analyzed is less than a MSE of the segment pair; in response to the MSE of the vector index being less than a worst MSE of the segment pair, updating the pair-wise initial candidate set to include the vector index; in response to a MSE of a current vector being less than a best MSE of the segment pair, updating the segment pair to include the vector index.
26. The encoder of any of claims 19-25, wherein, implementing a post-optimization on the plurality of pair-wise initial candidates includes: determining which pair-wise initial candidate of the plurality of pair-wise initial candidates has a lowest mean squared error (MSE) of the plurality of pair-wise initial candidates; for each pair-wise initial candidate other than the pair-wise initial candidate having the lowest MSE: comparing a MSE of the pair-wise initial candidate to a MSE of a neighbor vector in a forward direction; in response to the MSE of the pair-wise initial candidate being higher than the MSE of the neighbor vector in the forward direction, updating the pair-wise initial candidate to the neighbor vector in the forward direction; in response to the pair-wise initial candidate updated to the neighbor vector in the forward direction having a lower MSE than the pair-wise initial candidate having the lowest MSE, setting the pair-wise initial candidate updated to the neighbor vector in the forward direction as the pair-wise candidate having the lowest MSE; comparing a MSE of the pair-wise initial candidate to a MSE of a neighbor vector in a reverse direction; in response to the MSE of the pair-wise initial candidate being higher than the MSE of the neighbor vector in the reverse direction, updating the pair-wise initial candidate to the neighbor vector in the reverse direction; in response to the pair-wise initial candidate updated to the neighbor vector in the reverse direction having a lower MSE than the pair-wise initial candidate having the lowest MSE, setting the pair-wise initial candidate updated to the neighbor vector in the reverse direction as the pair-wise candidate having the lowest MSE.
27. The encoder of claim 26, further adapted to: comparing a MSE of the pair-wise initial candidate to a MSE of a next neighbor vector in the forward direction; updating the pair of initial candidates to a next neighbor vector in the positive direction in response to the MSE of the pair of initial candidates being higher than the MSE of the next neighbor vector in the positive direction; comparing the MSE of the pair of initial candidates to the MSE of a next neighbor vector in the negative direction; updating the pair of initial candidates to a next pair vector in the negative direction in response to the MSE of the pair of initial candidates being higher than the MSE of the next neighbor vector in the negative direction.
28. The encoder of claim 26 or 27, wherein, the neighbor vector and the next neighbor vector in the positive direction are part of a positive circular MSE neighbor index list and the neighbor vector and the next neighbor vector in the negative direction are part of a negative circular MSE neighbor index list.
29. The encoder of claim 28, wherein, the positive circular MSE neighbor index list and the negative circular MSE neighbor index list are created based on a vector ordering mse_order_circ across all the plurality of segments in a full cascade codebook cb_temporary_full with a length of Nv.
30. The encoder of any of claims 19-29, wherein, reconstructing the plurality of pair final candidates includes: obtaining an index idx_full in the range of [0…(Nv-1)] that contains a plurality of segments and corresponding column shift values in the range of 0 to Ns-1; performing a segment-wise and coefficient-wise upshift of the plurality of segments using the column shift values for the index to form upshifted segments; performing an inverse discrete cosine transform DCT Type II transform on the upshifted segments; scaling an output vector from the inverse DCT Type II transform to an original unscaled FDCNG domain vector; and add the global offset value vector back to the unscaled frequency domain comfort noise generation FD-CNG domain vector to form fdcng_final idx_full vector.
31. The encoder of claim 19, wherein, the truncated vector includes a vector segment with a different degree of high frequency content compared to other segments.
32. The encoder of claim 19, wherein, the final candidate data includes one or more of: a vector index, a reconstructed candidate vector, and / or a transformed codebook entry.
33. The encoder of claim 19, further adapted to perform a post-optimization on the plurality of pair initial candidates to replace a plurality of pair initial candidates of the plurality of pair initial candidates with candidate neighbor lists to form a plurality of final candidates.
34. A computer program comprising program code to be executed by a processing circuit (302) of an encoder (202), wherein, execution of the program code causes the encoder (202) to perform the operations of any of claims 1-15.
35. A computer program product comprising a non-transitory storage medium including program code to be executed by processing circuitry (302) of an encoder (202), wherein, execution of the program code causes the encoder (202) to perform the operations of any of claims 1-15.
36. An encoder (212) comprising: processing circuitry (402); and a memory (410) coupled with the processing circuitry, wherein the memory includes instructions that when executed by the processing circuitry cause the encoder to perform the operations of claim 16.
37. An encoder (212) adapted to perform the operations of claim 16.
38. A computer program comprising program code to be executed by a processing circuit (402) of a decoder (212), wherein, execution of the program code causes the encoder (212) to perform the operations of claim 16.
39. A computer program product comprising a non-transitory storage medium including program code to be executed by processing circuitry (402) of a decoder (212), wherein, execution of the program code causes the encoder (212) to perform the operations of claim 16. execution of the program code causes the encoder (212) to perform the operations of claim 16.