Audio code stream quality analysis and evaluation method based on cyanine color sound Audio Vivid

By decoding and threshold evaluation of Jingcaisheng Audio Vivid audio code streams, combined with objective and subjective evaluation, the problem that the existing technology cannot effectively evaluate the audio quality of three-dimensional sound technology is solved, and quantitative characterization and multi-scene monitoring support for Audio Vivid audio code stream quality are realized.

CN119943094AActive Publication Date: 2025-05-06HANGZHOU ARCVIDEO TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202311446443.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-05-06
Estimated Expiration
2043-11-02

AI Technical Summary

Technical Problem

The existing audio quality evaluation methods cannot effectively evaluate the dynamic metadata and the characteristics of sound beds, objects, and HOA sound field of complex objects in Jingcaisheng Audio Vivid, making it difficult for three-dimensional sound technology to make reasonable quality evaluations.

Method used

A method of audio code stream quality analysis and evaluation based on Jingcaisheng Audio Vivid is adopted. By obtaining the audio code stream and decoding, setting threshold indicators to evaluate the code stream and audio content information, combining objective and subjective evaluation, the quantitative characterization of the Audio Vivid audio code stream quality is realized.

Benefits of technology

It realizes quantitative characterization of Audio Vivid audio stream quality, supports multi-task evaluation, provides headphone and speaker monitoring solutions, can objectively evaluate based on specific scenarios, and supports mono and multi-channel monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943094A_ABST
    Figure CN119943094A_ABST
Patent Text Reader

Abstract

The invention discloses an audio code stream quality analysis and evaluation method based on audio Vivid, and the method comprises the steps: S1, obtaining an audio code stream of the audio Vivid, and inputting an IP code stream address of the audio code stream to a terminal; s2, the terminal transmits the audio code stream to a server, the server decodes the audio code stream and sends decoding information to the terminal for numerical representation, and audio objective quality evaluation is carried out on the audio code stream information; s3, the server encodes the decoded data, edits the decoded data into a compressed stream of the Flv-http protocol in an AAC format, pushes the compressed stream to a terminal for playing through websocket, supports dynamic switching of sound channels, and sets a corresponding monitoring scheme for earphone monitoring; s4, the server decodes the PCM data, connects the PCM data with an audio de-embedding device through an SDI line, connects a sound box according to a sound channel sequence for power amplification, and sets a corresponding monitoring scheme for sound box monitoring; and S5, performing comprehensive evaluation on the input audio code stream data through audio objective quality evaluation and audio subjective quality evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of audio processing, and in particular relates to an audio code stream quality analysis and evaluation method based on Audio Vivid. Background Art

[0002] Audio Vivid is a three-dimensional sound technology independently developed in China. Based on the audio of sound and objects, artists can create in three-dimensional space, including height, and the sound can be precisely placed and moved. It was jointly developed by the World Ultra High Definition Video Industry Alliance and the Digital Audio and Video Codec Technology Standard Working Group. By creating sound in three-dimensional space, the technology can break the channel limitation, give each sound a unique personality, and let the sound linger around and above the listener, achieving a stunning sense of reality.

[0003] At present, audio quality evaluation is mainly divided into two categories: subjective evaluation and objective evaluation. Subjective audio evaluation is performed by professionals who monitor the audio and evaluate the score according to their subjective feelings; however, since Audio Vivid has the characteristics of multi-channel sound bed, audio objects, and Ambisonic sound field compared to conventional audio, it is necessary to build a complex speaker space layout to clearly perceive the stereoscopic, surround, and spatial sense of three-dimensional sound during subjective evaluation, which has high requirements for the equipment environment. Objective audio evaluation methods usually include time domain analysis, frequency domain analysis, and acoustic feature analysis. Time domain analysis mainly analyzes the time series of audio signals, such as the waveform, amplitude, and time length of the signal; frequency domain analysis mainly analyzes the frequency components of audio signals, such as spectrum, frequency response, etc.; acoustic feature analysis mainly analyzes the sound characteristics of audio signals, such as timbre, volume, clarity, etc.; however, this solution cannot characterize the dynamic metadata of complex objects in Audio Vivid and describe and quantify the characteristics of sound bed, objects, and HOA (Ambisonic sound field).

[0004] In view of the above problems, the current 3D sound technology represented by AudioVivid cannot make reasonable quality evaluation from subjective and objective evaluation. Summary of the invention

[0005] In view of the above problems, the present invention provides an audio code stream quality analysis and evaluation method based on Audio Vivid.

[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0007] A method for analyzing and evaluating audio bitstream quality based on Audio Vivid, comprising the following steps:

[0008] S1, obtain the audio code stream of Audio Vivid, and input the IP code stream address of the audio code stream into the terminal;

[0009] S2, the terminal transmits the audio code stream to the server, the server decodes the audio code stream and sends the decoded information to the terminal for numerical representation, sets threshold indicators, sets indicators and trigger conditions related to the code stream and audio content information, deducts points based on a percentage system for each indicator, and when an event is triggered, calculates the deduction system for the indicator, and finally obtains a quantitative output evaluation, and performs an objective audio quality evaluation on the audio code stream information;

[0010] S3, the server encodes the decoded data and edits it into a compressed stream of the Flv-http protocol in the AAC format, pushes it to the terminal for playback through websocket, supports dynamic channel switching, sets the corresponding monitoring scheme for headphone monitoring, and implements subjective quality evaluation of headphone monitoring audio according to subjective evaluation indicators;

[0011] S4, the server decodes the PCM data and connects it to the audio de-embedder through the SDI line, connects the speakers to the amplifier according to the channel order, sets the corresponding monitoring scheme for speaker monitoring, and implements the subjective quality evaluation of the speaker monitoring audio according to the subjective evaluation index;

[0012] S5, through audio objective quality evaluation and audio subjective quality evaluation, the input audio bitstream data is comprehensively evaluated to obtain the Audio vivid audio bitstream quality evaluation result.

[0013] In a possible implementation, in S2, the decoded signal includes channel information, object signals, HOA signals, and corresponding metadata.

[0014] In a possible implementation, in S2, the numerical representation includes: source input audio type, source input audio bit depth, encoding method, sampling rate, channel name, channel input level, object channel number, object coordinate system, object current gain value, HOA layout, project information, channel sound field position, aggregate loudness value, loudness range value, maximum true peak, maximum instantaneous loudness and average dialogue loudness.

[0015] In a possible implementation, in S2, the objective audio quality evaluation includes the evaluation of the bitstream and the evaluation of the audio content information.

[0016] In a possible implementation, the code stream evaluation includes audio continuity error, packet loss, source status, and audio stream status.

[0017] In a possible implementation manner, the evaluation of the audio content information includes channel input level, loudness, channel signal, object signal, HOA signal and metadata.

[0018] In a possible implementation, the threshold indicators in S2 include bitstream-related threshold indicators and audio content information-related threshold indicators, wherein the bitstream-related threshold indicators include audio error duration, audio popping, audio silence packet loss duration, source status interruption time, and whether the audio stream status is lost. The corresponding duration is set as the trigger counting condition, and the number of times and deductions triggered from the start of monitoring to the end of monitoring are numerically weighted; the audio content information-related thresholds include the number of channels, whether the number of objects is fully represented, whether the HOA is represented, the channel input level, and the loudness trigger threshold, and the number of times and deductions triggered from the start of monitoring to the end of monitoring are numerically weighted.

[0019] In a possible implementation, the subjective evaluation indicators include sound quality, sense of space, sense of direction, sense of distance, sound source volume and sense of externality.

[0020] In a possible implementation, the terminal in S1 supports the input of audio code streams in Ts over UDP and Ts over http formats.

[0021] In a possible implementation, the monitoring scheme includes mono-channel monitoring, mono-channel polling monitoring and multi-channel polling monitoring.

[0022] The present invention has the following beneficial effects:

[0023] (1) The Audio Vivid audio bitstream quality analysis and evaluation method and business system are implemented, and the Audio Vivid audio bitstream is quantitatively characterized from both subjective and objective evaluation schemes.

[0024] (2) Using a task approach, it supports Audio Vivid audio stream quality analysis and evaluation for multiple input tasks without interfering with each other.

[0025] (3) In terms of subjective evaluation schemes, two schemes are proposed: monitoring through headphones and speakers. Objective evaluation can be carried out according to the schemes based on specific scenarios. At the same time, three monitoring schemes are supported: mono monitoring, mono polling monitoring, and multi-channel polling monitoring.

[0026] (4) In the objective evaluation scheme, conventional information is monitored through bitstream-related indicators, and Audio Vivid bitstream information is monitored through audio content information-related indicators. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 The present invention is a flowchart of the steps of the method for analyzing and evaluating the quality of an audio bitstream based on Audio Vivid. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0029] Reference Figure 1 , which is a flowchart of the step of the method for analyzing and evaluating the quality of an audio code stream based on Audio Vivid according to an embodiment of the present invention, comprising the following steps:

[0030] S1, obtain the Audio Vivid audio code stream, and input the IP code stream address of the audio code stream into the terminal; the terminal supports the input of audio code streams including Ts overUDP and Ts overhttp formats.

[0031] S2, the terminal transmits the audio code stream to the server, the server decodes the audio code stream and sends the decoded information to the terminal for numerical representation, sets threshold indicators, sets indicators and trigger conditions related to the code stream and audio content information, deducts points based on a percentage system for each indicator, and when an event is triggered, calculates the deduction system for the indicator, and finally obtains a quantitative output evaluation, and performs an objective audio quality evaluation on the audio code stream information;

[0032] S3, the server encodes the decoded data and edits it into a compressed stream of the Flv-http protocol in the AAC format, pushes it to the terminal for playback through websocket, supports dynamic channel switching, sets the corresponding monitoring scheme for headphone monitoring, and implements subjective quality evaluation of headphone monitoring audio according to subjective evaluation indicators;

[0033] S4, the server decodes the PCM data and connects it to the audio de-embedder through the SDI line, connects the speakers to the amplifier according to the channel order, sets the corresponding monitoring scheme for speaker monitoring, and implements the subjective quality evaluation of the speaker monitoring audio according to the subjective evaluation index;

[0034] S5, through audio objective quality evaluation and audio subjective quality evaluation, the input audio bitstream data is comprehensively evaluated to obtain the Audio vivid audio bitstream quality evaluation result.

[0035] In an embodiment of the present invention, an audio bitstream quality analysis and evaluation method based on Audio Vivid is provided. In S2, the decoded signal includes channel information, object signal, HOA signal and corresponding metadata. The numerical representation includes: source input audio type, source input audio bit depth, encoding method, sampling rate, channel name, channel input level, object channel number, object coordinate system, object current gain value, HOA layout, project information, channel sound field position, aggregate loudness value, loudness range value, maximum true peak value, maximum instantaneous loudness and average dialogue loudness. The objective audio quality evaluation includes the evaluation of the bitstream and the evaluation of the audio content information. The evaluation of the bitstream includes audio continuity errors, packet loss, source status and audio stream status. The evaluation of the audio content information includes channel input level, loudness, channel signal, object signal, HOA signal and metadata.

[0036] Furthermore, the threshold indicators in S2 include bitstream-related threshold indicators and audio content information-related threshold indicators, wherein the bitstream-related threshold indicators include audio error duration, audio popping, audio silence packet loss duration, source status interruption time, and whether the audio stream status is lost. The corresponding duration is set as the trigger counting condition, and the number of times the monitoring is triggered from the start of monitoring to the end of monitoring and the deduction points are numerically weighted; the audio content information-related thresholds include the number of channels, whether the number of objects is fully represented, whether the HOA is represented, the channel input level, and the loudness trigger threshold, and the number of times the monitoring is triggered from the start of monitoring to the end of monitoring and the deduction points are numerically weighted.

[0037] According to an embodiment of the present invention, the audio code stream quality analysis and evaluation method based on Audio Vivid is used, and the subjective evaluation indicators include sound quality, spatial sense, sense of direction, sense of distance, sound source volume and external sense. The sound quality includes clarity, balance, brightness, distortion, intimacy, and dynamics indicators; spatial sense includes reverberation perception and spatial perception indicators; sense of direction includes verticality and horizontality indicators; sense of distance includes loudness and sound pressure level indicators; sound source volume includes width and depth indicators; external sense includes indicators specifically for headphone monitoring. The configurable monitoring schemes include mono monitoring, mono polling monitoring, and multi-channel polling monitoring. Mono monitoring selects a single channel and drives the sound card for monitoring; mono polling monitoring selects multiple channels and sets the time interval for switching between channels to perform mono polling monitoring; multi-channel mixed monitoring selects multiple channels for mixed monitoring.

[0038] Through the above-mentioned setting, the Audio Vivid audio bitstream quality analysis and evaluation method based on Jingcaisheng Audio Vivid is realized, and the Audio Vivid audio bitstream quality analysis and evaluation method and business system are implemented, and the Audio Vivid audio bitstream is quantitatively characterized from two schemes: subjective evaluation and objective evaluation. In the objective evaluation scheme, conventional information is monitored through bitstream-related indicators, and Audio Vivid bitstream information is monitored through audio content information-related indicators. In the subjective evaluation scheme, two schemes of monitoring through headphones and speakers are proposed. Objective evaluation can be carried out according to the scheme according to the specific scenario. At the same time, it supports three monitoring scheme settings: mono monitoring, mono polling monitoring, and multi-channel polling monitoring.

[0039] It should be understood that the exemplary embodiments described herein are illustrative rather than restrictive. Although one or more embodiments of the present invention have been described in conjunction with the accompanying drawings, it should be understood by those skilled in the art that various changes in form and detail may be made without departing from the spirit and scope of the present invention as defined by the appended claims.

Claims

1. A method for analyzing and evaluating audio bitstream quality based on Audio Vivid, characterized in that: The following steps are involved: S1, obtain the audio code stream of Audio Vivid, and input the IP code stream address of the audio code stream into the terminal; S2, the terminal transmits the audio code stream to the server, the server decodes the audio code stream and sends the decoded information to the terminal for numerical representation, sets threshold indicators, sets indicators and trigger conditions related to the code stream and audio content information, deducts points based on a percentage system for each indicator, and when an event is triggered, calculates the deduction system for the indicator, and finally obtains a quantitative output evaluation, and performs an objective audio quality evaluation on the audio code stream information; S3, the server encodes the decoded data and edits it into a compressed stream of the Flv-http protocol in the AAC format, pushes it to the terminal for playback through websocket, supports dynamic channel switching, sets the corresponding monitoring scheme for headphone monitoring, and implements subjective quality evaluation of headphone monitoring audio according to subjective evaluation indicators; S4, the server decodes the PCM data and connects it to the audio de-embedder through the SDI line, connects the speakers to the amplifier according to the channel order, sets the corresponding monitoring scheme for speaker monitoring, and implements the subjective quality evaluation of the speaker monitoring audio according to the subjective evaluation index; S5, through audio objective quality evaluation and audio subjective quality evaluation, the input audio bitstream data is comprehensively evaluated to obtain the Audio vivid audio bitstream quality evaluation result.

2. The method for analyzing and evaluating the quality of an audio stream based on Audio Vivid according to claim 1, characterized in that: In S2, the decoded signal includes channel information, object signal, HOA signal and corresponding metadata.

3. The method for analyzing and evaluating the quality of an audio stream based on Audio Vivid according to claim 1, characterized in that: In S2, the numerical representations include: source input audio type, source input audio bit depth, encoding method, sampling rate, channel name, channel input level, object channel number, object coordinate system, object current gain value, HOA layout, project information, channel sound field position, aggregate loudness value, loudness range value, maximum true peak, maximum instantaneous loudness and average dialogue loudness.

4. The method for analyzing and evaluating the quality of an audio code stream based on Audio Vivid according to claim 1, characterized in that: In S2, the objective audio quality evaluation includes the evaluation of the bit stream and the evaluation of the audio content information.

5. The method for analyzing and evaluating the quality of an audio stream based on Audio Vivid as claimed in claim 4, characterized in that: The bitstream evaluation includes audio continuity errors, packet loss, source status, and audio stream status.

6. The method for analyzing and evaluating the quality of an audio stream based on Audio Vivid according to claim 4, characterized in that: The evaluation of audio content information includes channel input level, loudness, channel signal, object signal, HOA signal and metadata.

7. The method for analyzing and evaluating the quality of an audio stream based on Audio Vivid according to claim 1, characterized in that: The threshold indicators in S2 include bitstream-related threshold indicators and audio content information-related threshold indicators, where bitstream-related threshold indicators include audio error duration, audio popping, audio silence packet loss duration, source status interruption time, and whether the audio stream status is lost. The corresponding duration is set as the trigger counting condition, and the number of triggers and deductions from the start of monitoring to the end of monitoring are numerically weighted; The audio content information-related thresholds include the number of channels, whether the number of objects is fully represented, whether the HOA is represented, the channel input level, and the loudness trigger threshold. The number of triggers and deductions from the start of monitoring to the end of monitoring are numerically weighted.

8. The method for analyzing and evaluating the quality of an audio stream based on Audio Vivid according to claim 1, characterized in that: Subjective evaluation indicators include sound quality, sense of space, sense of direction, sense of distance, sound source volume and sense of externality.

9. The method for analyzing and evaluating the quality of an audio stream based on Audio Vivid according to any one of claims 1 to 8, characterized in that: The S1 terminal supports the input of audio code streams in Ts over UDP and Ts over http formats.

10. The method for analyzing and evaluating the quality of an audio stream based on Audio Vivid according to any one of claims 1 to 8, characterized in that: Monitoring solutions include mono monitoring, mono polling monitoring and multi-channel polling monitoring.

Citation Information

Patent Citations

  • Apparatus and method for providing measure of spatiality associated with audio stream

    CN110603820A

  • Audio and video multimedia database construction and multimedia subjective quality evaluation method

    CN111355949A

  • Audio quality recognition model training method and device, server and storage medium

    CN111863033A

  • Acoustic quality evaluation apparatus, acoustic quality evaluation method, and program

    US20220277765A1