Karaoke audio processing method, device and computer-readable storage medium

By obtaining the image and audio information of karaoke users, using neural network to determine user attributes, and automatically adjusting EQ sound effects processing, the problem of cumbersome adjustment of EQ sound effects by karaoke users is solved, and the karaoke experience is improved.

CN114974188BActive Publication Date: 2025-09-05BEIJING HORIZON INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210543147.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-18
Publication Date
2025-09-05
Estimated Expiration
2042-05-18

AI Technical Summary

Technical Problem

Ka Song users need to manually adjust the EQ sound effects, which is cumbersome to operate, which affects the Ka Song experience.

Method used

By obtaining image and/or audio information of karaoke users, the user attributes are determined using a neural network, and the EQ sound effect processing is automatically adjusted based on the user attributes.

Benefits of technology

It realizes that there is no need to manually adjust the EQ sound effects, and automatically adapts the corresponding EQ sound effects to different karaoke users to improve the karaoke experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114974188B_ABST
    Figure CN114974188B_ABST
Patent Text Reader

Abstract

Disclosed are a karaoke audio processing method, device, and computer-readable storage medium. The method comprises: obtaining image and / or audio information of a karaoke user; determining the karaoke user's attributes based on the image and / or audio information; and performing EQ sound processing on the karaoke user's vocal audio based on the user attributes. The disclosed embodiments can automatically adapt the corresponding EQ sound effects to different karaoke users, eliminating the need for manual EQ adjustment. This entire process is highly convenient and effectively ensures a superior karaoke experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to audio processing technology, and in particular to a karaoke audio processing method, device, and computer-readable storage medium. Background Art

[0002] Karaoke (for example, singing at a karaoke bar or using karaoke apps on mobile phones) is a popular pastime for many people. The EQ (equalizer) sound effect is a commonly used sound effect during karaoke. It can enhance the richness, brightness, and clarity of the music signal, increase the sense of presence, and highlight the timbre of instruments. However, in current karaoke scenarios, users need to manually adjust the EQ sound effect. Summary of the Invention

[0003] In order to solve the technical problem that users need to manually adjust the EQ sound effect, which is cumbersome to operate, the present disclosure is proposed. The embodiments of the present disclosure provide a karaoke audio processing method, device and computer-readable storage medium.

[0004] According to one aspect of the present disclosure, a karaoke audio processing method is provided, comprising:

[0005] Obtain image and / or audio information of karaoke users;

[0006] Determining user attributes of the karaoke user based on the image and / or audio information;

[0007] Based on the user attributes, EQ sound effect processing is performed on the vocal audio of the karaoke user.

[0008] According to another aspect of the present disclosure, a karaoke audio processing device is provided, comprising:

[0009] An acquisition module is used to acquire the image and / or audio information of the karaoke user;

[0010] A first determining module, configured to determine a user attribute of the karaoke user based on the image and / or audio information acquired by the acquiring module;

[0011] The first processing module is used to perform EQ sound effect processing on the vocal audio of the karaoke user based on the user attribute determined by the first determining module.

[0012] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and the computer program is used to execute the above-mentioned karaoke audio processing method.

[0013] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0014] processor;

[0015] a memory for storing instructions executable by the processor;

[0016] The processor is used to read the executable instructions from the memory and execute the instructions to implement the above-mentioned karaoke audio processing method.

[0017] Based on the karaoke audio processing method, device, computer-readable storage medium and electronic device provided by the above-mentioned embodiments of the present disclosure, the user attributes of the karaoke user can be determined based on the image and / or audio information of the karaoke user, and based on the determined user attributes, the EQ sound effect processing can be performed on the vocal audio of the karaoke user. In this way, the embodiments of the present disclosure can automatically adapt corresponding EQ sound effects to different karaoke users without the need to manually adjust the EQ sound effects. The whole process is very convenient to implement and can effectively ensure the karaoke experience.

[0018] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other purposes, features, and advantages of the present disclosure will become more apparent through a more detailed description of the embodiments of the present disclosure in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and are not intended to limit the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.

[0020] Figure 1 It is a flowchart of a karaoke audio processing method provided by an exemplary embodiment of the present disclosure.

[0021] Figure 2 It is a flowchart of a karaoke audio processing method provided by another exemplary embodiment of the present disclosure.

[0022] Figure 3 It is a flowchart of a karaoke audio processing method provided by another exemplary embodiment of the present disclosure.

[0023] Figure 4 It is a flowchart of a karaoke audio processing method provided by another exemplary embodiment of the present disclosure.

[0024] Figure 5 It is a flowchart of a karaoke audio processing method provided by another exemplary embodiment of the present disclosure.

[0025] Figure 6 It is a flowchart of a karaoke audio processing method provided by another exemplary embodiment of the present disclosure.

[0026] Figure 7It is a flowchart of a karaoke audio processing method provided by another exemplary embodiment of the present disclosure.

[0027] Figure 8 It is a flowchart of a karaoke audio processing method provided by another exemplary embodiment of the present disclosure.

[0028] Figure 9 2 is a schematic diagram of a target EQ adjustment method in an exemplary embodiment of the present disclosure.

[0029] Figure 10 FIG. 4 is a schematic structural diagram of a target EQ in an exemplary embodiment of the present disclosure.

[0030] Figure 11 FIG. 4 is a schematic structural diagram of a target EQ in another exemplary embodiment of the present disclosure.

[0031] Figure 12 It is a structural diagram of a karaoke audio processing device provided by an exemplary embodiment of the present disclosure.

[0032] Figure 13 It is a structural diagram of a karaoke audio processing device provided by another exemplary embodiment of the present disclosure.

[0033] Figure 14 3 is a structural diagram of a karaoke audio processing device provided by another exemplary embodiment of the present disclosure.

[0034] Figure 15 3 is a structural diagram of a karaoke audio processing device provided by another exemplary embodiment of the present disclosure.

[0035] Figure 16 3 is a structural diagram of a karaoke audio processing device provided by another exemplary embodiment of the present disclosure.

[0036] Figure 17 3 is a structural diagram of a karaoke audio processing device provided by another exemplary embodiment of the present disclosure.

[0037] Figure 18 3 is a structural diagram of a karaoke audio processing device provided by another exemplary embodiment of the present disclosure.

[0038] Figure 19 3 is a structural diagram of a karaoke audio processing device provided by another exemplary embodiment of the present disclosure.

[0039] Figure 20 is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0040] Below, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.

[0041] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.

[0042] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meanings, nor do they indicate a necessary logical order between them.

[0043] It should also be understood that in the embodiments of the present disclosure, “a plurality of” may refer to two or more than two, and “at least one” may refer to one, two, or more than two.

[0044] It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.

[0045] In addition, the term "and / or" in this disclosure is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this disclosure generally indicates that the related objects are in an "or" relationship.

[0046] It should also be understood that the description of the various embodiments in this disclosure focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced with each other. For the sake of brevity, they will not be described one by one.

[0047] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0048] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.

[0049] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0050] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0051] The embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate in conjunction with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, among others.

[0052] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media, including storage devices.

[0053] Application Overview

[0054] EQ sound effects can be used to improve the fullness, brightness and clarity of music signals, increase the sense of presence, and highlight the timbre of musical instruments. They are widely used in various scenarios such as karaoke, concerts, film production, and evening parties.

[0055] In the process of implementing the present disclosure, the inventors found that when EQ sound effects are applied to karaoke scenes, the EQ sound effects required by different karaoke users are often different. Therefore, the EQ sound effects need to be adjusted frequently in karaoke scenes. The current adjustment method of EQ sound effects is manual adjustment, which is very cumbersome to operate, thereby reducing the user's karaoke experience.

[0056] Exemplary Methods

[0057] Figure 1 It is a flowchart of a karaoke audio processing method provided by an exemplary embodiment of the present disclosure. Figure 1 The method shown includes step 110, step 120 and step 130, and each step is described below.

[0058] Step 110: Acquire the image and / or audio information of the karaoke user.

[0059] It should be noted that the karaoke user may be located in a target space where karaoke activities can be performed, such as in a car, a room, etc. In step 110, an image of the karaoke user may be captured by a camera provided in the target space to obtain image information of the karaoke user; and / or an audio signal of the karaoke user may be captured by a microphone provided in the target space to obtain audio information of the karaoke user.

[0060] Step 120: Determine the user attributes of the karaoke user based on the image and / or audio information.

[0061] In step 120, the karaoke user may be analyzed from at least one attribute dimension with reference to at least one of the image information and the audio information to determine a user attribute of the karaoke user. The user attribute of the karaoke user may include at least one attribute corresponding to the at least one attribute dimension. Optionally, the at least one attribute dimension includes, but is not limited to, an age dimension, a gender dimension, etc.

[0062] Step 130: Based on the user attributes, perform EQ sound effect processing on the vocal audio of the karaoke user.

[0063] During the karaoke process, the vocal audio of the karaoke user can be collected by a microphone set in the target space. Referring to the user attributes determined in step 120, the vocal audio can be processed with EQ sound effects to make the EQ sound effects of the vocal audio adapt to the karaoke user.

[0064] Based on the karaoke audio processing method provided by the above-mentioned embodiments of the present disclosure, the user attributes of the karaoke user can be determined based on the image and / or audio information of the karaoke user, and based on the determined user attributes, the EQ sound effect processing can be performed on the vocal audio of the karaoke user. In this way, the embodiments of the present disclosure can automatically adapt corresponding EQ sound effects to different karaoke users without the need to manually adjust the EQ sound effects. The whole process is very convenient to implement and can effectively ensure the karaoke experience.

[0065] In an optional example, the user attribute includes a first user attribute and a second user attribute.

[0066] Here, the first user attribute may be an age attribute, and the second user attribute may be a gender attribute.

[0067] It's important to note that people of different ages have different vocal ranges. Specifically, younger people's vocal organs are still developing, resulting in a narrower vocal range and limited singing ability. People in the middle age group have mature vocal organs, so their vocal range is relatively fixed. Older people's vocal organs gradually degenerate, resulting in a gradually narrowing vocal range. Due to these different vocal ranges, people of different ages require different EQ effects.

[0068] It's important to note that different genders have different vocal ranges. Specifically, men's vocal range is generally between G2-F2, while women's is generally between F3-D5. In other words, men's vocal range tends to be lower-frequency, while women's range tends to be higher-frequency. Due to these different vocal ranges, different genders require different EQ effects.

[0069] Since the vocal ranges of people of different age groups and people of different genders are different, in the embodiments of the present disclosure, the gender attributes and age attributes can be referred to to perform EQ sound effects processing on the vocal audio of the K song user, so that the EQ sound effects of the vocal audio can be adapted to the K song user, thereby improving the K song experience.

[0070] exist Figure 1 Based on the embodiment shown, Figure 2 As shown, step 120 includes step 1201 , step 1203 and step 1205 .

[0071] Step 1201: Acquire a face image of a karaoke user based on image information.

[0072] In step 1201, a face recognition algorithm may be used to extract a face image of the karaoke user from the image information.

[0073] Step 1203: extract features from the face image to obtain overall features and local features of the face.

[0074] In step 1203, an image feature extraction algorithm can be used to extract features from the facial image to obtain overall facial features and local facial features; wherein, overall facial features include overall features represented by brightness distribution information, such as the entire facial area image and skin color features of the facial image; local facial features include local features represented by the position and shape contour information of the facial features, such as the left eye image, right eye image, eye image, nose image, mouth image, wrinkle texture features, facial skull features, etc.

[0075] Step 1205 : Based on the overall facial features and local facial features, the user attributes of the karaoke user are determined via a pre-trained first neural network.

[0076] It should be noted that a large amount of sample data can be used in advance to train a first neural network for determining user attributes; each sample data may include input data and output data, the input data may include overall facial features and local facial features obtained by extracting features from a user's facial image, and the output data may include the user attributes of the user. Thus, in step 1205, the overall facial features and local facial features obtained in step 1203 need only be provided as input to the first neural network, which can then perform calculations based on the data to obtain the user attributes of the karaoke object.

[0077] In the embodiment of the present disclosure, by referring to the face image of the karaoke user and performing a feature extraction operation in combination with the use of the first neural network, the user attributes of the karaoke user can be determined efficiently and reliably.

[0078] exist Figure 1 Based on the embodiment shown, Figure 3 As shown, step 120 includes step 1207 , step 1209 and step 1211 .

[0079] Step 1207: Based on the audio information, obtain the target audio of the karaoke user.

[0080] It should be noted that when the audio of the karaoke user is collected through the microphone, some interference signals may be collected at the same time. In view of this, in step 1207, these interference signals can be filtered out from the audio signal to obtain the target audio of the karaoke user.

[0081] Step 1209: extract features from the object audio to obtain audio features.

[0082] In step 1209, an audio feature extraction algorithm may be used to extract features from the target audio to obtain audio features. The audio features include but are not limited to Mel-frequency cepstral coefficients, fundamental pitch, harmonics, and other features.

[0083] Step 1211: Based on the audio features, determine the user attributes of the karaoke user via a pre-trained second neural network.

[0084] It should be noted that a second neural network for determining user attributes can be trained in advance using a large amount of sample data; each sample data may include input data and output data, the input data may include audio features extracted from a user's audio object, and the output data may include the user attributes of the user. Thus, in step 1211, the audio features obtained in step 1209 need only be provided as input to the second neural network, which can then perform calculations based on the audio features to obtain the user attributes of the karaoke object.

[0085] In the embodiment of the present disclosure, by referring to the audio features obtained by extracting the object audio of the karaoke user and combining it with the use of the second neural network, the user attributes of the karaoke user can be determined efficiently and reliably.

[0086] exist Figure 1 Based on the embodiment shown, Figure 4 As shown, step 120 includes step 1213 , step 1215 , step 1217 and step 1219 .

[0087] Step 1213: Based on the image information and audio information, obtain the face image and object audio of the karaoke user.

[0088] Step 1215: extract features from the facial image to obtain overall facial features and local facial features.

[0089] Step 1217: extract features from the object audio to obtain audio features.

[0090] It should be noted that the specific implementation process of step 1213 can refer to the description of step 1201 and step 1207, the specific implementation process of step 1215 can refer to the description of step 1203, and the specific implementation process of step 1217 can refer to the description of step 1209, which will not be repeated here.

[0091] Step 1219: Determine the user attributes of the karaoke user based on the overall facial features, local facial features, and audio features.

[0092] In step 1219, the user attributes of the karaoke user can be determined via a first neural network based on the overall facial features and local facial features, and the user attributes of the karaoke user can be determined via a second neural network based on the audio features. Then, the user attributes determined by the first neural network and the second neural network are combined to determine the final user attributes of the karaoke user.

[0093] Of course, in a specific implementation, a large amount of sample data can also be used in advance to train a third neural network for determining user attributes; wherein each sample data may include input data and output data, the input data may include overall facial features and local facial features obtained by extracting features from a user's facial image, and audio features obtained by extracting features from the user's object audio, and the output data may include the user attributes of the user. Thus, in step 1219, the overall facial features and local facial features obtained in step 1215, as well as the audio features obtained in step 1217, can be provided as input to the third neural network, and the third neural network can perform operations based on the data to obtain the user attributes of the karaoke object.

[0094] In the embodiments of the present disclosure, by combining the face image and the object audio of the karaoke user and performing a feature extraction operation, the user attributes of the karaoke user can be determined efficiently and reliably.

[0095] In the case where the user attributes include a first user attribute and a second user attribute, Figure 1 Based on the embodiment shown, Figure 5 As shown, step 130 includes step 1301 and step 1303 .

[0096] Step 1301 , when the first user attribute is the first preset attribute or the second preset attribute, EQ sound effect processing is performed on the vocal audio of the karaoke user according to the first preset rule or the second preset rule.

[0097] Here, the first user attribute may be an age attribute, the first preset attribute may be a low age group, the second preset attribute may be a high age group, the first preset rule may be an EQ sound effect processing rule adapted for the low age group, and the second preset rule may be an EQ sound effect processing rule adapted for the high age group.

[0098] Step 1303: When the first user attribute is the third preset attribute, EQ sound effect processing is performed on the vocal audio of the karaoke user based on the second user attribute.

[0099] Here, the third preset attribute may be the middle age group.

[0100] In the embodiments of the present disclosure, for people in the younger age group, EQ sound effect processing can be performed on their vocal audio according to the EQ sound effect processing rules suitable for the younger age group. For people in the older age group, EQ sound effect processing can be performed on their vocal audio according to the EQ sound effect processing rules suitable for the older age group. For people in the middle age group, EQ sound effect processing can be performed on their vocal audio with reference to their gender, so that the EQ sound effect of the vocal audio can be effectively adapted to the K song users, thereby improving the K song experience.

[0101] exist Figure 5 Based on the embodiment shown, Figure 6 As shown, the method further includes step 140 , step 150 and step 160 .

[0102] Step 140 : When the first user attribute is the first preset attribute, the volume of the human voice audio after EQ sound effect processing is increased.

[0103] In step 140, if the age attribute of the karaoke user is in the young age group (the karaoke user can be considered a child), the volume of the vocal audio after EQ sound effect processing can be automatically increased by a preset value, such as 10, 15, 20, etc.

[0104] Step 150: Mix the original vocals and accompaniment of the karaoke user's karaoke song and the vocal audio to obtain a mixed audio.

[0105] It should be noted that the server can store the original vocals and accompaniments of multiple songs. In step 150, the name of the current K song program of the K song user can be used as index information to search in the server to obtain the original vocals and accompaniment of the K song program, so as to mix the obtained original vocals and accompaniment with the human voice audio.

[0106] Step 160: Play the mixed audio through an audio playback device.

[0107] Optionally, the audio playback device may include a speaker.

[0108] Generally speaking, children's vocal organs are still developing. If they are not used properly and scientifically during karaoke, it will affect the development of the vocal cords, and in severe cases, it may even lead to loss of voice. In addition, children's voices are short and shallow, with little power and a narrow range. In view of this, in the embodiments of the present disclosure, when the karaoke user is a child, the volume of the human voice audio can be automatically increased, and the original singer and accompaniment can be turned on. This can not only protect the child's vocal cords, but also allow the child to sing along, thereby ensuring the child's karaoke experience.

[0109] exist Figure 5 Based on the embodiment shown, Figure 7 As shown, the method further includes step 170 , step 180 and step 190 .

[0110] Step 170 : When the first user attribute is the first preset attribute, the karaoke posture of the karaoke object is determined based on the image information.

[0111] In step 170, for the case where the age attribute of the karaoke user is in the young age group (the karaoke user can be considered as a child), a posture detection algorithm can be used to perform posture detection on the image information of the karaoke object, so as to determine the posture of the karaoke object when singing karaoke, and this posture is the karaoke posture.

[0112] Step 180: Outputting karaoke posture correction prompt information based on the karaoke posture satisfying the preset correction conditions.

[0113] In step 180, it can be determined whether the karaoke posture meets the preset correction conditions. For example, if the karaoke subject's posture during karaoke is hunched or the head is tilted, it can be determined that the karaoke posture meets the preset correction conditions. At this time, the karaoke mode can be exited, for example, by stopping the playback of the mixed audio and stopping the collection of human voice audio. In addition, karaoke posture correction prompt information can be output in the form of voice, text, etc. to prompt the karaoke user to keep muscles relaxed, back straight, head upright, etc. After the karaoke posture correction prompt information is output, the karaoke mode can be re-entered, for example, by continuing to play the mixed audio and continue to collect human voice audio.

[0114] Of course, when the karaoke posture meets the preset correction conditions, it is also possible not to exit the karaoke mode, but only output the karaoke posture correction prompt information.

[0115] Step 190: exit the karaoke mode based on the karaoke time of the karaoke object exceeding the preset time.

[0116] In the case where the karaoke object is a child, its continuous karaoke time can be detected. If the detected continuous karaoke time exceeds the preset time, such as more than 15 minutes, 20 minutes, etc., the karaoke mode can be exited. At this time, the microphone used by the child can be turned off, and the child will not be able to continue karaoke. In addition, rest reminder information can be output in the form of voice, text, etc. to remind the child to take a rest.

[0117] In the embodiment of the present disclosure, if the karaoke user is a child, karaoke posture correction prompt information can be output to prompt the child to adopt the correct posture for karaoke, and when the child sings for a long time, the karaoke mode can be automatically exited to allow the child to rest, thereby protecting the child's vocal cords.

[0118] exist Figure 1 Based on the embodiment shown, Figure 8 As shown, step 130 includes step 1305 , step 1307 and step 1309 .

[0119] Step 1305: Determine the target EQ adjustment method based on the user attributes.

[0120] It should be noted that the target EQ is a device used to implement EQ sound effect processing, and the adjustment methods of the target EQ include but are not limited to increasing the low-frequency gain, decreasing the low-frequency gain, increasing the high-frequency gain, and decreasing the high-frequency gain.

[0121] In the case where the user attributes include a first user attribute and a second user attribute, the first user attribute is an age attribute and the second user attribute is a gender attribute, such as Figure 9 As shown, there can be the following rules:

[0122] Rule 1: If the age attribute is low, the target EQ is adjusted by increasing the low-frequency gain and the high-frequency gain;

[0123] Rule 2: If the age attribute is high, the target EQ is adjusted by increasing the low-frequency gain and the high-frequency gain;

[0124] Rule 3: If the age attribute is middle age and the gender attribute is male, the target EQ adjustment method is: increase the high frequency gain;

[0125] Rule 4: If the age attribute is middle age and the gender attribute is female, the target EQ adjustment method is: increase the low-frequency gain.

[0126] It should be noted that the above rule 1 can be used as at least part of the first preset rule above, for example, Figure 9 The first preset rule may also include rules such as turning up the volume and turning on the original singer; the above-mentioned rule 2 may serve as at least part of the second preset rule mentioned above.

[0127] Step 1307: Adjust the target EQ according to the adjustment method.

[0128] In a specific embodiment, the target EQ includes N filters with different center frequencies, and as Figure 10 As shown, the N filters are N bandpass filters (i.e., bandpass filter 1 to bandpass filter N) arranged in parallel, or, as shown Figure 11 As shown, the N filters are a low shelving filter, N-2 peaking filters (i.e., peaking filter 1 to peaking filter N-2) and a high shelving filter arranged in series;

[0129] Step 1307 includes:

[0130] When the adjustment mode is to increase the low-frequency gain, the gain of at least one filter having a low-frequency center frequency among the N filters is increased;

[0131] When the adjustment mode is to reduce the low-frequency gain, reducing the gain of at least one filter whose center frequency belongs to the low frequency among the N filters;

[0132] When the adjustment mode is to increase the high frequency gain, the gain of at least one filter having a center frequency belonging to the high frequency among the N filters is increased;

[0133] When the adjustment method is to reduce the high frequency gain, the gain of at least one filter having a high frequency center frequency among the N filters is reduced.

[0134] Optionally, N can be 10, 15, 27, 31 or other values, which are not listed here one by one.

[0135] In an example, the target EQ includes 10 filters (i.e., N is 10), and the center frequencies of these 10 filters are: 31.5Hz, 63Hz, 125Hz, 250Hz, 500Hz, 1000Hz, 2000Hz, 4000Hz, 8000Hz, and 16000Hz, respectively. Among them, 31.5Hz, 63Hz, 125Hz, and 250Hz are low frequencies, 500Hz, 1000Hz, and 2000Hz are mid-frequencies, and 4000Hz, 8000Hz, and 16000Hz are high frequencies. In this way, when the adjustment method is to increase the low-frequency gain, the gain of at least one of the four filters corresponding to 31.5Hz, 63Hz, 125Hz, and 250Hz can be increased; when the adjustment method is to lower the low-frequency gain, the gain of at least one of the four filters corresponding to 31.5Hz, 63Hz, 125Hz, and 250Hz can be reduced; when the adjustment method is to increase the high-frequency gain, the gain of at least one of the four filters corresponding to 4000Hz, 8000Hz, and 16000Hz can be increased; when the adjustment method is to lower the high-frequency gain, the gain of at least one of the four filters corresponding to 4000Hz, 8000Hz, and 16000Hz can be reduced.

[0136] Step 1309 : Perform EQ sound effect processing on the human voice audio based on the adjusted target EQ.

[0137] After adjusting the target EQ according to the adjustment method, the human voice audio can be provided to the adjusted target EQ frame by frame. Each filter in the adjusted target EQ will be processed accordingly based on its own gain (for example, after multiplying with the gain, the multiplication results will be added), thereby ultimately realizing the EQ sound effect processing of the human voice audio.

[0138] In the embodiments of the present disclosure, referring to the user attributes of the karaoke user, the adjustment method of the target EQ can be reasonably determined, and the target EQ is subsequently adjusted based on the adjustment method, and the human voice audio is processed with EQ sound effects based on the adjusted target EQ, which can effectively ensure that the EQ sound effects of the human voice audio are adapted to the karaoke user, thereby ensuring the karaoke experience.

[0139] It should be emphasized that in the embodiments of the present disclosure, the acquisition, storage and application of user attribute related information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0140] Any karaoke audio processing method provided in the embodiments of the present disclosure can be executed by any appropriate device with data processing capabilities, including but not limited to: a terminal device and a server. Alternatively, any karaoke audio processing method provided in the embodiments of the present disclosure can be executed by a processor, such as a processor that executes any karaoke audio processing method mentioned in the embodiments of the present disclosure by calling corresponding instructions stored in a memory. This will not be further described below.

[0141] Exemplary devices

[0142] Figure 12 It is a structural diagram of a karaoke audio processing device provided by an exemplary embodiment of the present disclosure. Figure 12 The device shown includes an acquisition module 1210 , a first determination module 1220 and a first processing module 1230 .

[0143] Acquisition module 1210, for acquiring image and / or audio information of karaoke users;

[0144] A first determining module 1220 is configured to determine a user attribute of a karaoke user based on the image and / or audio information acquired by the acquiring module 1210;

[0145] The first processing module 1230 is used to perform EQ sound effect processing on the vocal audio of the karaoke user based on the user attributes determined by the first determining module 1220.

[0146] In an optional example, the user attribute includes a first user attribute and a second user attribute.

[0147] In an alternative example, Figure 13 As shown, the first determining module 1220 includes:

[0148] The first acquisition submodule 12201 is used to acquire a face image of a karaoke user based on the image information acquired by the acquisition module 1210;

[0149] The second acquisition submodule 12203 is used to extract features from the facial image acquired by the first acquisition submodule 12201 to obtain overall facial features and local facial features;

[0150] The first determination submodule 12205 is used to determine the user attributes of the karaoke user through a pre-trained first neural network based on the overall facial features and local facial features obtained by the second acquisition submodule 12203.

[0151] In an alternative example, Figure 14 As shown, the first determining module 1220 includes:

[0152] The third acquisition submodule 12207 is used to acquire the target audio of the karaoke user based on the audio information acquired by the acquisition module 1210;

[0153] The fourth acquisition submodule 12209 is configured to extract features from the object audio acquired by the third acquisition submodule 12207 to obtain audio features.

[0154] The second determining submodule 12211 is used to determine the user attributes of the karaoke user through a pre-trained second neural network based on the audio features obtained by the fourth obtaining submodule 12209.

[0155] In an alternative example, Figure 15 As shown, the first determining module 1220 includes:

[0156] The fifth acquisition submodule 12213 is used to acquire the face image and audio of the karaoke user based on the image information and audio information acquired by the acquisition module 1210;

[0157] The sixth acquisition submodule 12215 is configured to perform feature extraction on the facial image acquired by the fifth acquisition submodule 12213 to obtain overall facial features and local facial features.

[0158] a seventh acquisition submodule 12217, configured to extract features from the object audio acquired by the fifth acquisition submodule 12213 to obtain audio features;

[0159] The third determination submodule 12219 is used to determine the user attributes of the karaoke user based on the overall facial features and local facial features obtained by the sixth acquisition submodule 12215 and the audio features obtained by the seventh acquisition submodule 12217.

[0160] In an alternative example, Figure 16 As shown, the first processing module 1230 includes:

[0161] The first processing submodule 12301 is configured to perform EQ sound effect processing on the vocal audio of the karaoke user according to the first preset rule or the second preset rule when the first user attribute is the first preset attribute or the second preset attribute;

[0162] The second processing submodule 12303 is configured to perform EQ sound effect processing on the vocal audio of the karaoke user based on the second user attribute when the first user attribute is the third preset attribute.

[0163] In an alternative example, Figure 17 As shown, the device also includes:

[0164] The volume adjustment module 1240 is configured to increase the volume of the human voice audio after the EQ sound effect processing by the first processing module 1230 when the first user attribute is the first preset attribute;

[0165] The mixing module 1250 is used to mix the original singing and accompaniment of the karaoke user's karaoke song and the human voice audio to obtain a mixed audio;

[0166] The playing module 1260 is configured to play the mixed audio obtained by the mixing module 1250 through an audio playing device.

[0167] In an alternative example, Figure 18 As shown, the device also includes:

[0168] The second determining module 1270 is configured to determine the karaoke posture of the karaoke object based on the image information acquired by the acquiring module 1210 when the first user attribute is the first preset attribute;

[0169] The second processing module 1280 is configured to output a karaoke posture correction prompt message based on the karaoke posture determined by the second determining module 1270 to meet a preset correction condition;

[0170] The exit module 1290 is used to exit the karaoke mode if the karaoke time of the karaoke object exceeds a preset time.

[0171] In an alternative example, Figure 19 As shown, the first processing module 1230 includes:

[0172] The fourth determining submodule 12305 is configured to determine an adjustment method for the target EQ based on the user attributes determined by the first determining module 1220;

[0173] an adjusting submodule 12307, configured to adjust the target EQ according to the adjustment method determined by the fourth determining submodule 12305;

[0174] The third processing submodule 12309 is used to perform EQ sound effect processing on the human voice audio based on the target EQ adjusted by the adjustment submodule 12307.

[0175] In an optional example, the target EQ includes N filters with different center frequencies, and the N filters are N bandpass filters arranged in parallel, or the N filters are a low-shelving filter, N-2 peak filters, and a high-shelving filter arranged in series;

[0176] The adjustment submodule 12307 includes:

[0177] a first adjustment unit, configured to increase the gain of at least one filter having a low-frequency center frequency among the N filters when the adjustment mode determined by the fourth determination submodule is to increase the low-frequency gain;

[0178] a second adjustment unit, configured to reduce the gain of at least one filter having a low-frequency center frequency among the N filters when the adjustment mode determined by the fourth determination submodule is to reduce the low-frequency gain;

[0179] a third adjustment unit, configured to increase the gain of at least one filter having a high frequency center frequency among the N filters when the adjustment mode determined by the fourth determination submodule is to increase the high frequency gain;

[0180] The fourth adjustment unit is configured to reduce the gain of at least one filter having a high frequency center frequency among the N filters when the adjustment mode determined by the fourth determining submodule is to reduce the high frequency gain.

[0181] Exemplary electronic devices

[0182] Below, reference Figure 20 The electronic device according to the embodiment of the present disclosure is described. The electronic device may be either or both of the first device and the second device, or a standalone device independent of them, and the standalone device may communicate with the first device and the second device to receive collected input signals from them.

[0183] Figure 20 A block diagram of an electronic device according to an embodiment of the present disclosure is illustrated.

[0184] like Figure 20 As shown, the electronic device 2000 includes one or more processors 2010 and a memory 2020 .

[0185] The processor 2010 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 2000 to perform desired functions.

[0186] The memory 2020 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 2010 may run the program instructions to implement the karaoke audio processing method of each embodiment of the present disclosure described above and / or other desired functions. Various contents such as audio data may also be stored in the computer-readable storage medium.

[0187] In one example, the electronic device 2000 may further include an input device 2030 and an output device 2040 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0188] For example, when the electronic device 2000 is the first device or the second device, the input device 2030 may be a microphone or a microphone array to input received human voice data. When the electronic device 2000 is a standalone device, the input device 2030 may be a communication network connector to receive collected input signals from the first device and the second device.

[0189] In addition, the input device 2030 may also include, for example, a keyboard, a mouse, etc. The input device 2030 may be used to input song request information.

[0190] The output device 2040 can output various information to the outside, including the mixed audio mentioned above, etc. The output device 2040 can include, for example, a display, a speaker, a printer, a communication network and its connected remote output device, etc.

[0191] Of course, to simplify, Figure 20 Only some of the components related to the present disclosure in the electronic device 2000 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 2000 may further include any other appropriate components.

[0192] Exemplary computer program products and computer-readable storage media

[0193] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the karaoke audio processing method according to various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section of this specification.

[0194] The computer program product may be written in any combination of one or more programming languages ​​to implement the operations of the disclosed embodiments, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0195] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enables the processor to execute the steps of the karaoke audio processing method according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.

[0196] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0197] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.

[0198] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments are sufficient. For system embodiments, since they largely correspond to method embodiments, their description is relatively simple. For relevant parts, references to the description of the method embodiments are sufficient.

[0199] The block diagrams of the devices, devices, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0200] The methods and apparatus of the present disclosure may be implemented in many ways. For example, the methods and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the methods according to the present disclosure. Thus, the present disclosure also covers recording media that store programs for executing the methods according to the present disclosure.

[0201] It should also be noted that in the apparatus, device, and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.

[0202] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0203] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A karaoke audio processing method, comprising: Obtain image and / or audio information of karaoke users; Determining user attributes of the karaoke user based on the image and / or audio information; Based on the user attributes, EQ sound effect processing is performed on the vocal audio of the karaoke user; wherein the user attributes are used to determine whether to increase the volume of the vocal audio after the EQ sound effect processing; The user attributes include a first user attribute and a second user attribute, and performing EQ sound effect processing on the vocal audio of the karaoke user based on the user attributes includes: When the first user attribute is the first preset attribute or the second preset attribute, performing EQ sound effect processing on the vocal audio of the karaoke user according to the first preset rule or the second preset rule; When the first user attribute is the third preset attribute, EQ sound effect processing is performed on the vocal audio of the karaoke user based on the second user attribute.

2. The method according to claim 1, wherein The determining of the user attributes of the karaoke user based on the image and / or audio information includes: Based on the image information, obtaining a facial image of the karaoke user; Performing feature extraction on the facial image to obtain overall facial features and local facial features; Based on the overall facial features and the local facial features, the user attributes of the karaoke user are determined via a pre-trained first neural network.

3. The method according to claim 1, wherein Determining the user attributes of the karaoke user based on the image and / or audio information includes: Based on the audio information, obtaining the target audio of the karaoke user; Performing feature extraction on the object audio to obtain audio features; Based on the audio features, user attributes of the karaoke user are determined via a pre-trained second neural network.

4. The method according to claim 1, wherein Determining the user attributes of the karaoke user based on the image and / or audio information includes: Based on the image information and the audio information, obtaining a face image and object audio of the karaoke user; Performing feature extraction on the facial image to obtain overall facial features and local facial features; Performing feature extraction on the object audio to obtain audio features; The user attributes of the karaoke user are determined based on the overall facial features, the local facial features and the audio features.

5. The method according to claim 1, further comprising: When the first user attribute is the first preset attribute, increasing the volume of the human voice audio after the EQ sound effect processing; Mixing the original singing and accompaniment of the karaoke song by the karaoke user and the vocal audio to obtain a mixed audio; Play the mixed audio through an audio playback device.

6. The method according to claim 1, further comprising: In a case where the first user attribute is the first preset attribute, determining a karaoke posture of the karaoke user based on the image information; Outputting karaoke posture correction prompt information based on the karaoke posture meeting the preset correction conditions; The karaoke mode is exited based on the karaoke user's karaoke time exceeding a preset time.

7. The method according to any one of claims 1 to 6, wherein: The performing EQ sound effect processing on the vocal audio of the karaoke user based on the user attributes includes: Determining a target EQ adjustment method based on the user attributes; Adjusting the target EQ according to the adjustment method; Perform EQ sound effect processing on the vocal audio based on the adjusted target EQ.

8. The method according to claim 7, wherein: The target EQ includes N filters with different center frequencies, and the N filters are N bandpass filters arranged in parallel, or the N filters are a low-shelving filter, N-2 peak filters, and a high-shelving filter arranged in series; The adjusting the target EQ according to the adjustment method includes: When the adjustment method is to increase the low-frequency gain, increasing the gain of at least one filter whose center frequency belongs to the low frequency among the N filters; When the adjustment mode is to lower the low-frequency gain, reducing the gain of at least one filter whose center frequency belongs to the low frequency among the N filters; When the adjustment method is to increase the high-frequency gain, increasing the gain of at least one filter having a high-frequency center frequency among the N filters; When the adjustment method is to lower the high-frequency gain, the gain of at least one filter having a high-frequency center frequency among the N filters is reduced.

9. A karaoke audio processing device, comprising: An acquisition module is used to acquire the image and / or audio information of the karaoke user; A first determining module, configured to determine a user attribute of the karaoke user based on the image and / or audio information acquired by the acquiring module; A first processing module, configured to perform EQ sound effect processing on the vocal audio of the karaoke user based on the user attribute determined by the first determining module; wherein the user attribute is used to determine whether to increase the volume of the vocal audio after the EQ sound effect processing; The user attributes include a first user attribute and a second user attribute, and the first processing module includes: A first processing submodule is configured to perform EQ sound effect processing on the vocal audio of the karaoke user according to a first preset rule or a second preset rule when the first user attribute is a first preset attribute or a second preset attribute; The second processing submodule is used to perform EQ sound effect processing on the vocal audio of the karaoke user based on the second user attribute when the first user attribute is the third preset attribute.

10. A computer-readable storage medium storing a computer program, wherein the computer program is used to execute the karaoke audio processing method according to any one of claims 1 to 8.

11. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the karaoke audio processing method described in any one of claims 1-8.

Citation Information

Patent Citations

  • A music recommendation method and a mobile terminal

    CN106407424A

  • 10-segment parameter equalizer

    CN110971213A

  • Microphone parameter setting system

    CN111898537A

  • Karaoke machine and game machine

    JP2007304619A

  • Audio playback method and apparatus, and loudspeaker device

    WO2019174081A1