Audio vivid audio stream quality analysis and evaluation method

By setting threshold indicators and designing a monitoring scheme for the Audio Vivid audio stream, the problem that existing technologies cannot effectively evaluate the audio quality of Audio Vivid was solved, and a comprehensive quality analysis and evaluation of its audio stream was achieved.

CN119943094BActive Publication Date: 2025-11-21HANGZHOU ARCVIDEO TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311446443.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-11-21
Estimated Expiration
2043-11-02

AI Technical Summary

Technical Problem

Existing audio quality assessment methods cannot effectively evaluate the complex object dynamic metadata and soundbed, object, and HOA characteristics of Audio Vivid 3D sound technology, resulting in the inability to make reasonable quality assessments.

Method used

The audio stream quality analysis and evaluation method based on Audio Vivid is adopted. By setting threshold indicators and trigger conditions, a deduction system is used to calculate the audio stream and content information. Combined with headphone and speaker monitoring, a comprehensive analysis of subjective and objective evaluation is achieved.

Benefits of technology

It enables quantitative characterization of Audio Vivid audio stream quality, supports multi-task interference-free analysis, provides mono and multi-channel monitoring solutions, and achieves comprehensive quality evaluation of Audio Vivid audio streams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943094B_ABST
    Figure CN119943094B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on cyan color sound Audio Vivid audio stream quality analysis and evaluation method, comprising: S1, obtain cyan color sound Audio Vivid audio stream, and the IP stream address of audio stream is input to terminal;S2, terminal audio stream is transmitted to server, server decodes audio stream and sends decoding information to terminal for numerical characterization, audio objective quality evaluation is carried out to audio stream information;S3, server decodes data are encoded, and edited into the compression stream of Flv-http protocol of AAC format, by websocket push is played to terminal, support dynamic switching sound channel, set corresponding monitoring scheme and carry out earphone monitoring;S4, server decodes PCM data are connected audio demosaic by SDI line, according to sound channel order connection speaker is power amplifier, set corresponding monitoring scheme and carry out speaker monitoring;S5, by audio objective quality evaluation and audio subjective quality evaluation, the input audio stream data is comprehensively evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of audio processing technology, specifically relating to a method for quality analysis and evaluation of audio bitstreams based on Audio Vivid. Background Technology

[0002] Audio Vivid is a 3D sound technology independently developed in China. Based on sound and object audio, it allows artists to create in three-dimensional space, including height, with precise placement and movement of sound. It was jointly developed by the World Ultra HD Video Industry Alliance and the Digital Audio and Video Coding Technology Standards Working Group. By creating sound in three-dimensional space, this technology breaks through channel limitations, giving each sound a unique personality and allowing sound to resonate around and above the listener, achieving a stunning sense of realism.

[0003] Current audio quality assessment is mainly divided into two categories: subjective evaluation and objective evaluation. Subjective audio evaluation involves professionals listening to the audio and assigning scores based on their subjective impressions. However, due to the multi-channel sound bed, audio objects, and ambisonic sound field characteristics of Audio Vivid compared to conventional audio, subjective evaluation requires a complex speaker layout to clearly perceive the stereo, surround, and spatial effects of three-dimensional sound, placing high demands on the equipment environment. Objective audio evaluation methods typically include time-domain analysis, frequency-domain analysis, and acoustic feature analysis. Time-domain analysis mainly analyzes the time series of the audio signal, such as waveform, amplitude, and duration; frequency-domain analysis mainly analyzes the frequency components of the audio signal, such as the spectrum and frequency response; acoustic feature analysis mainly analyzes the sonic characteristics of the audio signal, such as timbre, volume, and clarity. However, this approach cannot characterize the dynamic metadata of complex objects in Audio Vivid or describe and quantify the characteristics of the sound bed, objects, and HOA (Ambisonic Sound Field).

[0004] Given the aforementioned problems, current 3D sound technology, represented by AudioVivid, is unable to make a reasonable quality assessment from both subjective and objective perspectives. Summary of the Invention

[0005] In view of the above-mentioned problems, the present invention provides a method for quality analysis and evaluation of audio bitstream based on Audio Vivid.

[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0007] A method for quality analysis and evaluation of audio bitstreams based on Audio Vivid includes the following steps:

[0008] S1, obtain the Audio Vivid audio stream and input the IP address of the audio stream into the terminal;

[0009] S2, the terminal transmits the audio bitstream to the server. The server decodes the audio bitstream and sends the decoded information to the terminal for numerical representation. By setting threshold indicators, indicators and triggering conditions related to bitstream and audio content information are set. Each indicator is scored out of 100. When an event is triggered, the indicators are calculated using the deduction system, and finally a quantitative output evaluation is obtained to make an objective audio quality evaluation of the audio bitstream information.

[0010] S3: The server encodes the decoded data and edits it into a compressed stream of AAC format using the Flv-http protocol. This stream is then pushed to the terminal for playback via WebSocket. It supports dynamic channel switching, setting up corresponding monitoring schemes for headphone monitoring, and achieving subjective evaluation of headphone audio quality according to subjective evaluation indicators.

[0011] S4, the server decodes PCM data and connects it to the audio de-embedding unit via SDI cable. The speakers are then connected to the amplifier in the order of the channels. The corresponding monitoring scheme is set up for speaker monitoring. Subjective evaluation of the audio quality of the speakers is achieved according to subjective evaluation indicators.

[0012] S5 performs a comprehensive evaluation of the input audio bitstream data through objective and subjective audio quality assessments to obtain the Audio Vivid audio bitstream quality evaluation result.

[0013] In one possible implementation, in S2, the decoded signal includes channel information, object signal, HOA signal, and corresponding metadata.

[0014] In one possible implementation, in S2, the numerical representation includes: source input audio type, source input audio bit depth, encoding method, sampling rate, channel name, channel input level, object channel number, object coordinate system, object current gain value, HOA layout, item information, channel sound field position, aggregate loudness value, loudness range value, maximum true peak value, maximum instantaneous loudness, and average dialogue loudness.

[0015] In one possible implementation, in S2, the objective audio quality evaluation includes the evaluation of the bitstream and the evaluation of audio content information.

[0016] In one possible implementation, the evaluation of the bitstream includes audio continuity errors, packet loss, source status, and audio stream status.

[0017] In one possible implementation, the evaluation of audio content information includes channel input level, loudness, channel signal, object signal, HOA signal, and metadata.

[0018] In one possible implementation, the threshold indicators in S2 include bitstream-related threshold indicators and audio content information-related threshold indicators. The bitstream-related threshold indicators include the duration of audio errors, audio pops, duration of audio silence packet loss, source status interruption time, and whether the audio stream status is lost. The corresponding duration is set as the trigger counting condition, and the number of triggers and deductions from the start of monitoring to the end of monitoring are numerically weighted. The audio content information-related threshold indicators include the number of channels, whether the number of objects is fully represented, whether the HOA is represented, channel input level, and loudness trigger threshold. The number of triggers and deductions from the start of monitoring to the end of monitoring are numerically weighted.

[0019] In one possible implementation, subjective evaluation indicators include sound quality, spatial sense, directionality, distance sense, sound source volume, and external feel.

[0020] In one possible implementation, the terminal in S1 supports input of audio streams in TS over UDP and TS over HTTP formats.

[0021] One possible implementation includes a monitoring scheme that includes mono monitoring, mono polling monitoring, and multi-channel polling monitoring.

[0022] The present invention has the following beneficial effects:

[0023] (1) A method and business system for analyzing and evaluating the quality of Audio Vivid audio streams were implemented, and the Audio Vivid audio streams were quantitatively characterized from two schemes: subjective evaluation and objective evaluation.

[0024] (2) The task-based approach supports the analysis and evaluation of the audio Vivid audio bitstream quality for multiple input tasks, with no interference between tasks.

[0025] (3) Regarding the subjective evaluation scheme, two schemes are proposed: one using headphones and the other using speakers. Objective evaluation can be carried out according to the specific scenario. At the same time, three monitoring schemes are supported: mono monitoring, mono polling monitoring, and multi-channel polling monitoring.

[0026] (4) In terms of objective evaluation, routine information monitoring is carried out through bitstream-related indicators, and audio content-related indicators are used to monitor Audio Vivid bitstream information. Attached Figure Description

[0027] Figure 1 This is a flowchart illustrating the steps of the Audio Vivid audio stream quality analysis and evaluation method based on an embodiment of the present invention. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] Reference Figure 1 The diagram shows a flowchart of the steps in the Audio Vivid audio stream quality analysis and evaluation method according to an embodiment of the present invention, including the following steps:

[0030] S1, obtain the Audio Vivid audio stream and input the IP address of the audio stream to the terminal; the terminal supports input of audio streams in TS over UDP and TS over HTTP formats.

[0031] S2, the terminal transmits the audio bitstream to the server. The server decodes the audio bitstream and sends the decoded information to the terminal for numerical representation. By setting threshold indicators, indicators and triggering conditions related to bitstream and audio content information are set. Each indicator is scored out of 100. When an event is triggered, the indicators are calculated using the deduction system, and finally a quantitative output evaluation is obtained to make an objective audio quality evaluation of the audio bitstream information.

[0032] S3: The server encodes the decoded data and edits it into a compressed stream of AAC format using the Flv-http protocol. This stream is then pushed to the terminal for playback via WebSocket. It supports dynamic channel switching, setting up corresponding monitoring schemes for headphone monitoring, and achieving subjective evaluation of headphone audio quality according to subjective evaluation indicators.

[0033] S4, the server decodes PCM data and connects it to the audio de-embedding unit via SDI cable. The speakers are then connected to the amplifier in the order of the channels. The corresponding monitoring scheme is set up for speaker monitoring. Subjective evaluation of the audio quality of the speakers is achieved according to subjective evaluation indicators.

[0034] S5 performs a comprehensive evaluation of the input audio bitstream data through objective and subjective audio quality assessments to obtain the Audio Vivid audio bitstream quality evaluation result.

[0035] An embodiment of the present invention provides a method for quality analysis and evaluation of audio bitstreams based on Audio Vivid. In step S2, the decoded signal includes channel information, object signal, HOA signal, and corresponding metadata. Numerical representations include: source input audio type, source input audio bit depth, encoding method, sampling rate, channel name, channel input level, object channel number, object coordinate system, object current gain value, HOA layout, item information, channel sound field position, aggregate loudness value, loudness range value, maximum true peak value, maximum instantaneous loudness, and average dialogue loudness. Objective audio quality evaluation includes evaluation of the bitstream and evaluation of audio content information. Bitstream evaluation includes audio continuity errors, packet loss, source status, and audio stream status. Audio content information evaluation includes channel input level, loudness, channel signal, object signal, HOA signal, and metadata.

[0036] Furthermore, the threshold indicators in S2 include bitstream-related threshold indicators and audio content information-related threshold indicators. The bitstream-related threshold indicators include the duration of audio errors, audio pops, duration of audio silence packet loss, source status interruption time, and whether the audio stream status is lost. The corresponding duration is set as the trigger counting condition, and the number of triggers and deductions from the start of monitoring to the end of monitoring are numerically weighted. The audio content information-related threshold indicators include the number of channels, whether the number of objects is fully represented, whether the HOA is represented, channel input level, and loudness trigger threshold. The number of triggers and deductions from the start of monitoring to the end of monitoring are numerically weighted.

[0037] An embodiment of this invention provides a method for quality analysis and evaluation of audio streams based on Audio Vivid. Subjective evaluation indicators include sound quality, spatial awareness, directionality, distance perception, sound source volume, and externalization perception. Sound quality includes indicators such as clarity, balance, brightness, distortion, intimacy, and dynamics; spatial awareness includes reverberation perception and spatial perception indicators; directionality includes verticality and horizontality indicators; distance perception includes indicators such as loudness and sound pressure level; sound source volume includes width and depth indicators; and externalization perception includes indicators specifically for headphone monitoring. Configurable monitoring schemes include mono monitoring, mono polling monitoring, and multi-channel polling monitoring. Mono monitoring involves selecting a single channel and driving the sound card for monitoring; mono polling monitoring involves selecting multiple channels and setting the time interval between channel switching; and multi-channel mixed monitoring involves selecting multiple channels for mixed monitoring.

[0038] The above-described implementation of the Audio Vivid audio stream quality analysis and evaluation method based on Jingcaisheng has resulted in a comprehensive Audio Vivid audio stream quality analysis and evaluation system. It provides a quantitative characterization of the Audio Vivid audio stream through both subjective and objective evaluation methods. For the objective evaluation, routine information monitoring is performed using stream-related metrics, while audio content-related metrics are used to monitor Audio Vivid stream information. For the subjective evaluation, two monitoring schemes are proposed: headphone and speaker monitoring. Objective evaluation can be performed according to the specific scenario, and three monitoring schemes are supported: mono monitoring, mono polling monitoring, and multi-channel polling monitoring.

[0039] It should be understood that the exemplary embodiments described herein are illustrative and not restrictive. Although one or more embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of the invention as defined by the appended claims.

Claims

1. A method for quality analysis and evaluation of audio bitstreams based on Audio Vivid, characterized in that, Includes the following steps: S1, obtain the Audio Vivid audio stream and input the IP address of the audio stream into the terminal; S2, the terminal transmits the audio bitstream to the server. The server decodes the audio bitstream and sends the decoded information to the terminal for numerical representation. By setting threshold indicators, indicators and triggering conditions related to bitstream and audio content information are set. Each indicator is scored out of 100. When an event is triggered, the indicators are calculated using the deduction system, and finally a quantitative output evaluation is obtained to make an objective audio quality evaluation of the audio bitstream information. S3: The server encodes the decoded data and edits it into a compressed stream of AAC format using the Flv-http protocol. This stream is then pushed to the terminal for playback via WebSocket. It supports dynamic channel switching, setting up corresponding monitoring schemes for headphone monitoring, and achieving subjective evaluation of headphone audio quality according to subjective evaluation indicators. S4, the server decodes PCM data and connects it to the audio de-embedding unit via SDI cable. The speakers are then connected to the amplifier in the order of the channels. The corresponding monitoring scheme is set up for speaker monitoring. Subjective evaluation of the audio quality of the speakers is achieved according to subjective evaluation indicators. S5 performs a comprehensive evaluation of the input audio bitstream data through objective and subjective audio quality assessments to obtain the Audio Vivid audio bitstream quality evaluation result.

2. The method for quality analysis and evaluation of audio bitstream based on Audio Vivid as described in claim 1, characterized in that, In S2, the decoded signal includes channel information, object signal, HOA signal, and corresponding metadata.

3. The method for quality analysis and evaluation of audio bitstream based on Audio Vivid as described in claim 1, characterized in that, In S2, the numerical representation includes: source input audio type, source input audio bit depth, encoding method, sampling rate, channel name, channel input level, object channel number, object coordinate system, object current gain value, HOA layout, item information, channel sound field position, aggregate loudness value, loudness range value, maximum true peak value, maximum instantaneous loudness, and average dialogue loudness.

4. The method for quality analysis and evaluation of audio bitstream based on Audio Vivid as described in claim 1, characterized in that, In S2, the objective audio quality evaluation includes the evaluation of the bitstream and the evaluation of audio content information.

5. The method for quality analysis and evaluation of audio bitstream based on Audio Vivid as described in claim 4, characterized in that, The evaluation of the bitstream includes audio continuity errors, packet loss, source status, and audio stream status.

6. The method for quality analysis and evaluation of audio bitstream based on Audio Vivid as described in claim 4, characterized in that, The evaluation of audio content information includes channel input level, loudness, channel signal, object signal, HOA signal, and metadata.

7. The method for quality analysis and evaluation of audio bitstream based on Audio Vivid as described in claim 1, characterized in that, The threshold indicators in S2 include bitstream-related threshold indicators and audio content information-related threshold indicators. Bitstream-related threshold indicators include audio error duration, audio popping, audio mute packet loss duration, source status interruption time, and whether the audio stream status is lost. The corresponding duration is set as the trigger counting condition, and the number of times the trigger is performed and the deduction is numerically weighted from the start of monitoring to the end of monitoring. When considering audio content information related thresholds, including the number of channels, whether the number of objects is fully represented, whether HOA is represented, channel input level, and loudness trigger threshold, the number of triggers and deductions from the start of monitoring to the end of monitoring are numerically weighted.

8. The method for quality analysis and evaluation of audio bitstream based on Audio Vivid as described in claim 1, characterized in that, Subjective evaluation indicators include sound quality, spatial sense, direction sense, distance sense, sound source volume, and external feel.

9. The method for quality analysis and evaluation of audio bitstream based on Audio Vivid as described in any one of claims 1 to 8, characterized in that, The S1 terminal supports input of audio streams in TS over UDP and TS over HTTP formats.

10. The method for quality analysis and evaluation of audio bitstream based on Audio Vivid as described in any one of claims 1 to 8, characterized in that, The monitoring solutions include mono monitoring, mono polling monitoring, and multi-channel polling monitoring.

Citation Information

Patent Citations

  • Audio and video multimedia database construction and multimedia subjective quality evaluation method

    CN111355949A

  • Acoustic quality evaluation apparatus, acoustic quality evaluation method, and program

    US20220277765A1