Karaoke method, device, electronic equipment and readable storage medium

By acquiring and recognizing lip area images through intelligent display devices, and collecting and adjusting karaoke audio data, the problems of poor effect and high cost in home karaoke are solved, and the karaoke effect is improved without the use of external devices.

CN115631737BActive Publication Date: 2026-02-27SHENZHEN SKYWORTH RGB ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210981474.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-15
Publication Date
2026-02-27
Estimated Expiration
2042-08-15

AI Technical Summary

Technical Problem

Home karaoke systems struggle to balance karaoke quality and cost, especially when external karaoke devices are not used, resulting in poor performance and increased expenses.

Method used

The system acquires images of the target user's lip area through a smart display device, performs image recognition to determine the singing state, collects karaoke audio data, adjusts audio frequencies to enhance audio data from the target user's direction and reduce audio data from non-target user directions, and enables audio data playback.

Benefits of technology

Without relying on external karaoke equipment, the system enhances the vocal proportions of the target user, achieving the desired karaoke effect and reducing karaoke costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115631737B_ABST
    Figure CN115631737B_ABST
Patent Text Reader

Abstract

The application discloses a karaoke method and device, electronic equipment and a readable storage medium, and is applied to the technical field of audio processing. The karaoke method comprises the following steps: acquiring a lip region image of a target user in a karaoke mode; determining whether the target user is in a singing state by performing image recognition on the lip region image; if yes, collecting karaoke audio data of the target user in the karaoke mode, wherein the karaoke audio data comprises first audio data in the direction of the target user and second audio data in the direction of non-target users; adjusting the first audio data by using a first preset audio frequency point and adjusting the second audio data by using a second preset audio frequency point to obtain to-be-played audio data; and playing the to-be-played audio data so that the target user and an intelligent display device can perform karaoke interaction according to the to-be-played audio data. The application solves the technical problem that it is difficult to balance the karaoke effect and the karaoke cost in home karaoke.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio processing, in particular to a karaoke method and device, electronic equipment and readable storage medium. BACKGROUND

[0002] With the continuous development of science and technology, intelligent devices play an increasingly important role in people's lives. In some home application scenarios, the interaction between intelligent devices and external devices is required to meet the user's use demand, such as karaoke and games. At present, the family karaoke usually needs to increase the volume of the human voice through a karaoke microphone or other karaoke external device to achieve the expected karaoke effect. However, this method relies too much on karaoke equipment. If the user is limited by subjective and objective factors and lacks the corresponding karaoke equipment, the expected effect cannot be achieved when karaoke is performed. Therefore, the current family karaoke cannot balance the karaoke effect and the cost of karaoke. SUMMARY

[0003] The main purpose of the present application is to provide a karaoke method, device, electronic equipment and readable storage medium, which aims to solve the technical problem that the existing family karaoke cannot balance the karaoke effect and the cost of karaoke.

[0004] To achieve the above-mentioned purpose, the present application provides a karaoke method applied to an intelligent display device, which comprises:

[0005] acquiring a lip region image of a target user in a karaoke mode;

[0006] determining whether the target user is in a singing state by image recognition on the lip region image;

[0007] acquiring a lip region image of a target user in a karaoke mode;

[0008] determining whether the target user is in a singing state by image recognition on the lip region image;

[0009] If yes, collecting karaoke audio data corresponding to the target user in the karaoke mode, wherein the karaoke audio data includes first audio data in the direction of the target user and second audio data in the direction of non-target user;

[0010] adjusting the first audio data by a first preset audio frequency point and adjusting the second audio data by a second preset audio frequency point to obtain to-be-played audio data, wherein the first preset audio frequency point is greater than the second preset audio frequency point;

[0011] playing the to-be-played audio data for the target user and the intelligent display device to perform karaoke interaction according to the to-be-played audio data.

[0012] To achieve the above object, the application further provides a karaoke device applied to a smart display device, the karaoke device comprising:

[0013] an acquisition module configured to acquire a lip region image of a target user in a karaoke mode;

[0014] a determination module configured to determine whether the target user is in a singing state by performing image recognition on the lip region image;

[0015] an acquisition module configured to acquire, if yes, karaoke audio data corresponding to the target user in the karaoke mode, wherein the karaoke audio data comprises first audio data in a target user direction and second audio data in a non-target user direction;

[0016] an adjustment module configured to adjust the first audio data by a first preset audio frequency point and to adjust the second audio data by a second preset audio frequency point to obtain to-be-played audio data, wherein the first preset audio frequency point is greater than the second preset audio frequency point;

[0017] a playing module configured to play the to-be-played audio data so that the target user and the smart display device perform karaoke interaction according to the to-be-played audio data.

[0018] The application further provides an electronic device, which comprises a memory, a processor, and a program of the karaoke method stored in the memory and executable on the processor, and the program of the karaoke method can realize the steps of the karaoke method when executed by the processor.

[0019] The application further provides a computer readable storage medium, which stores a program of a karaoke method, and the program of the karaoke method can realize the steps of the karaoke method when executed by a processor.

[0020] The application further provides a computer program product, which comprises a computer program, and the computer program can realize the steps of the karaoke method when executed by a processor.

[0021] The application provides a karaoke method and device, electronic equipment and readable storage medium, which are applied to a smart display device, that is, a lip region image of a target user is acquired in a karaoke mode; whether the target user is in a singing state is determined by image recognition on the lip region image, that is, whether the user is in the karaoke is determined by image recognition on the lip region image of the target user; then if yes, the corresponding karaoke audio data of the target user is collected in the karaoke mode, wherein the karaoke audio data includes first audio data in the direction of the target user and second audio data in the direction of non-target user; then the first audio data is adjusted by a first preset audio frequency point, and the second audio data is adjusted by a second preset audio frequency point to obtain to-be-played audio data, wherein the first preset audio frequency point is greater than the second preset audio frequency point; and the to-be-played audio data is played for the target user and the smart display device to interact in the karaoke according to the to-be-played audio data. When the target user is in the karaoke mode, the audio frequency point corresponding to the audio data in the direction of the target user is consistent with the audio frequency point corresponding to the audio data in the direction of non-target user, then the first audio data in the direction of the target user is enhanced by the first preset audio frequency point, and the second audio data in the direction of non-target user is weakened by the second preset audio frequency point, so that the purpose of enhancing the audio data in the target direction is achieved, that is, the karaoke voice ratio of the target user is directly amplified, and the expected karaoke effect is achieved, that is, even if the user does not purchase a karaoke device, the karaoke voice ratio can be amplified when the user is in the karaoke, and when the user is in the karaoke mode and is singing, the karaoke voice ratio is indirectly amplified by a microphone or other karaoke external device, so that the technical defect that the karaoke effect is poor due to not purchasing a karaoke device considering the karaoke cost is overcome, and the purpose of considering the karaoke effect and the karaoke cost is achieved. BRIEF DESCRIPTION OF DRAWINGS

[0022] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the application.

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows, and obviously, other drawings can be obtained by those skilled in the art without any creative labor under the premise of the drawings.

[0024] Figure 1 The flowchart of the first embodiment of the karaoke method of the present application;

[0025] Figure 2 The family karaoke scene diagram of the karaoke method of the present application;

[0026] Figure 3 a flowchart of a KTV method according to a second embodiment of the present application;

[0027] Figure 4 a schematic diagram of a KTV device according to an embodiment of the present application;

[0028] Figure 5 a device structure schematic diagram of a hardware running environment involved in a KTV method according to an embodiment of the present application.

[0029] The purposes, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0030] In order to make the above-mentioned purposes, features and advantages of the present application more apparent and easy to understand, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings of the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0031] Embodiment one

[0032] First of all, it should be understood that, with the influence of subjective and objective factors, family KTV has become the mainstream choice of users, among which, family TV KTV is favored by people due to its simple operation and strong experience, etc. At present, if a user wants to perform family KTV on a smart TV, he usually needs to use a KTV microphone and other KTV external devices to complete the amplification of the vocal proportion, so as to achieve the expected KTV effect. However, since the KTV external device is only used for KTV, its cost performance is not high, and many users do not want to increase their KTV cost, so if a family user performs KTV without purchasing a KTV external device, the KTV effect will not be satisfactory. For example, even if the user tries to shorten the relative distance between the smart TV, the smart TV cannot output the sound effect that the user is satisfied with after picking up the user's KTV audio. Therefore, in the current family KTV scene, there is a technical pain point that the KTV effect is poor without using KTV external devices, and using KTV devices will increase the KTV cost. Therefore, there is an urgent need for a method that takes into account the KTV effect and KTV cost.

[0033] The present application provides a KTV method applied to a smart display device. In a first embodiment of the KTV method of the present application, referring to Figure 1 , the KTV method comprises:

[0034] Step S10: acquiring a lip region image of a target user in a KTV mode;

[0035] Step S20, determining whether the target user is in a singing state by image recognition on the lip region image;

[0036] In the embodiment, it is to be noted that the K-song method is applied to a smart display device, which is provided with a camera and a microphone array, such as a smart television, etc., and the K-song mode is triggered by a K-song system according to a K-song instruction, which can be a K-song instruction input by a user or automatically triggered by the user at a preset time point. The lip region image is used to represent an image of a lip region of the target user, which is intercepted by the camera of the smart display device and used to determine whether the target user is singing, and can be an image composed of a lip region intercepted from a user face image of the target user or a lip region image of the target user directly intercepted, for example, in an implementable manner, assuming that the smart display device is a smart television, and a user A pre-sets an online K-song function of a K-song application program to be started at 19:50 every Friday, when the online K-song function is triggered, the smart television is in a K-song mode, at this time, the camera of the smart television captures a user face frame, and extracts the lip region image from the user face frame, wherein the specific manner of extracting the lip region image can be to locate the lips based on a face feature extraction model and a pattern matching algorithm, to confirm a center point and an inner and outer contour, to further determine contour points according to an adaptive algorithm, and finally to generate a lip region image consistent with parameters of the user face frame according to the contour points.

[0037] In addition, it is to be noted that the target user is a unique user determined by the K-song system to sing in the K-song mode, which can be determined by the camera by recognizing a user in a detection region, for example, in an implementable manner, assuming that there is only one face in a user face frame captured by the camera, the face corresponding to the user is the target user, if there are multiple faces in the user face frame, the target user can be determined according to other facial feature information, for example, assuming that there are user B and user C face images in the user face image, the eye features of user B and user C can be obtained, and the user looking at the smart display device is determined as the target user, if user B and user C are both looking at the smart display device, the posture feature information of user B and user C can be further used to determine the target user.

[0038] Additionally, it should be noted that, since the user is the target user, the target user's attention is focused on karaoke for a certain period of time, and further, by identifying the opening and closing state of the target user's lips, it can be further determined whether the target user is singing, and the singing state is used to represent that the target user is karaoke, for example, in an implementable manner, by pre-storing a preset closed image of the target user's lips in a closed state, and then by comparing whether the lip region image and the preset closed image are consistent, if the lip region image and the preset closed image are inconsistent, it can be considered that the target user is in a singing state.

[0039] As an example, steps S10 to S20 include: capturing a user face image in a karaoke mode through a camera, if there is a single face feature in the user face image, the user corresponding to the face feature is taken as a target user, and the lip region feature of the target user is extracted from the user face image, if there are multiple face features in the user face image, a target face feature is determined from a specific face feature among the face features, and the user corresponding to the target face feature is taken as the target user, and the lip region feature of the target user is extracted from the user face image, wherein the specific feature can be an eye feature or a body feature, etc.; comparing the lip region image with a preset closed image for image recognition, if the lip image feature and the preset closed image are consistent, it is determined that the target user is not in the singing state, if the lip image and the preset closed image are inconsistent, it is determined that the target user is in the singing state, wherein the comparison method can be a lip image key point comparison method.

[0040] In an implementable manner, with reference to Figure 2 , Figure 2 To represent a home karaoke scene diagram, wherein the smart display device 13 includes a camera 13 and a microphone array 14, and the target user 11 emits a human voice audio, wherein the camera 13 is used to capture a user face image of the target user 11, and when the target user 11 emits a human voice audio, it can be collected by the microphone array 14.

[0041] Wherein the lip region image includes a first lip key frame image and a second lip key frame image, and the step of determining whether the target user is in a singing state by image recognition of the lip region image includes:

[0042] Step A10, based on a preset image feature extraction model, respectively extracting image features from the first lip key frame image and the second lip key frame image to obtain first image features and second image features;

[0043] Step A20, based on a preset image feature extraction model, image feature extraction is performed on the first lip key frame image and the second lip key frame image respectively to obtain first image features and second image features.

[0044] In this embodiment, it should be noted that a single lip region image cannot reflect the dynamic change of the target user's lips in a certain time period, for example, it is assumed that the user happens to open his mouth at the current time period, but it is possible that he is just yawning, so the dynamic change of the target user's lips can be accurately reflected through multiple lip region images, and the first lip key frame image and the second lip key frame image are continuous lip region images, wherein the first lip key frame image is used to represent the lip region image corresponding to the mouth region of the current face frame, and the second lip key frame image is used to represent the lip region image corresponding to the mouth region of the previous frame of the current face frame, and the first lip key frame image and the second lip key frame image are both key frames of an image sequence, and the preset image feature extraction model is a pre-trained image feature extraction model.

[0045] As an example, steps A10 to A20 include: inputting the first lip key frame image and the second lip key frame image into a preset image feature extraction model respectively to extract features of the entire first lip key frame image and the entire second lip key frame image, to obtain first image features corresponding to the first local region image and second image features corresponding to the second local region image; comparing the similarity of the first image features and the second image features to obtain a similarity comparison result corresponding to the first image features, and determining whether the target user is in a singing state according to the similarity comparison result, wherein the similarity comparison result includes similarity and dissimilarity, when the similarity comparison result is similarity, it is determined that the target user is not in the singing state, otherwise, it is in the singing state.

[0046] Step S30, if yes, the KTV audio data corresponding to the target user is collected in the KTV mode, wherein the KTV audio data includes first audio data in the direction of the target user and second audio data in the direction of non-target users;

[0047] Step S40, adjusting the first audio data by a first preset audio frequency point and adjusting the second audio data by a second preset audio frequency point to obtain to-be-played audio data, wherein the first preset audio frequency point is greater than the second preset audio frequency point;

[0048] Step S50, playing the to-be-played audio data for the target user and the intelligent display device to perform KTV interaction according to the to-be-played audio data.

[0049] In this embodiment, it should be noted that the K-song audio data is used to represent the audio in each direction collected by the microphone array of the smart display device, when it is determined that the target user is singing in the K-song mode, the far-field human voice part in the direction of the target user is amplified, while the interference sound in other directions is suppressed, where the interference sound can be the audio output by other users or environmental noise, and then the audio data in the direction of the target user (first audio data) and the audio data in the direction of the target user (second audio data) are determined after collecting the K-song audio data of the target user, so that the desired K-song effect can be achieved without additional K-song equipment, therefore, the first preset audio frequency point is used to enhance the audio frequency point corresponding to the first audio data, and the second preset audio frequency point is used to weaken the audio frequency point corresponding to the second audio data, where the desired K-song effect is at least consistent with the effect of singing by the user through the K-song external device.

[0050] As an example, steps S30 to S50 include: if it is determined that the target user is in the singing state, collecting K-song audio data corresponding to the target user in the K-song mode through the microphone array of the smart display device, where the K-song audio data includes first audio data in the direction of the target user and second audio data in the direction of the target user; adjusting the first audio data through a first preset audio frequency point and adjusting the second audio data through a second preset audio frequency point to obtain to-be-played audio data, where the first preset audio frequency point is greater than the second preset audio frequency point; playing the target K-song audio through the microphone array of the smart display device for K-song interaction between the target user and the smart display device according to the to-be-played audio data, where the K-song interaction is an interaction in the K-song mode.

[0051] The step of collecting K-song audio data corresponding to the target user in the K-song mode includes:

[0052] Step B10, collecting song external audio data in the K-song mode, where the song external audio data includes human voice audio data and interference audio data;

[0053] Step B20, if the human voice audio data is single-person voice audio data, the single-person voice audio data is taken as the first audio data in the direction of the target user, and the interference audio data is taken as the second audio data in the direction of the target user;

[0054] Step B30, if the human voice audio data is multi-human voice audio data, determining first audio data of the target user direction in the multi-human voice audio data;

[0055] Step B40, determining second audio data of the non-user direction according to the first audio data and the interference audio data.

[0056] In this embodiment, it should be noted that the song external audio data is used to represent audio obtained outside the smart display device, which can be environmental sound or human voice, etc., wherein the human voice audio obtained outside the smart display device is the human voice audio data, and other audio data obtained outside the smart display device is collectively referred to as the interference audio data, the multi-human voice audio data is used to represent human voice audio data in which there is target user human voice audio data and other user human voice audio data, the single-human voice audio data is used to represent human voice audio data that only has the target user human voice audio data, and when the human voice audio data is multi-human voice audio data, the non-user direction second audio data includes interference audio data and other user human voice audio data.

[0057] As an example, steps B10 to B40 include: collecting song external audio data through a microphone array of the smart display device in the K-song mode, performing segmentation processing on the K-song audio to obtain the human voice audio data and the interference audio data; extracting human voice print features in a preset human voice frequency range, if the human voice print features are single-human voice print features, determining that the human voice audio data is the single-human voice audio data, and then taking the single-human voice audio data as the first audio data of the target user direction and taking the interference audio data as the second audio data of the non-user direction, wherein the preset human voice frequency range can be 82Hz-392Hz; if the human voice print features are multi-human voice print features, determining first audio data of the target user direction in the multi-human voice audio data; taking audio data in the multi-human voice audio data other than the first audio data and the interference audio data together as the second audio data of the non-user direction.

[0058] Before the step of adjusting the first audio data through the first preset audio frequency point and adjusting the second audio data through the second preset audio frequency point to obtain the to-be-played audio data, the K-song method further includes:

[0059] Step C10, collecting song audio data and historical K-song audio data of the target user in the K-song mode;

[0060] Step C20, obtaining the corresponding relationship between the user audio frequency point corresponding to the historical K song audio data and the song audio frequency point corresponding to the song audio data;

[0061] Step C30, determining the first preset audio frequency point according to the corresponding relationship and the song audio data corresponding to the song audio data.

[0062] In this embodiment, it should be noted that the song audio data is collected by the microphone array of the smart display device, and the song audio data is used to represent the original song audio of the K song system at the current time point obtained from the cloud of the smart display device or the preset storage area. The historical K song audio data is used to represent the historical K song level of the target user, which can be embodied in the extreme pitch value of the target user in the past K song mode, including historical K song low frequency and historical K song high frequency. For example, in an implementable manner, assuming that the lowest frequency point of the original song audio is x and the highest frequency point is y, the average singing situation of the target user in the last three times of singing the lowest frequency point x and the highest frequency point y is x1, x2, x3 and y1, y2, y3, respectively. The historical K song low frequency point is the average value of x1, x2, x3, and the historical K song high frequency point is the average value of y1, y2, y3.

[0063] In addition, it should be noted that when the human voice audio data is the multi-person voice audio data, it means that the microphone array of the smart display device collects not only the human voice audio of the target user, but also the human voice audio of other users. The audio type can be talking audio or singing audio, etc. In addition to the audio data in the direction of the target user, the audio data in other directions is interference audio data. The target user direction represents the direction of the target user, and the first audio frequency point and the second audio frequency point represent the specific frequency point of the audio. The specific frequency point can be the lowest frequency point and the highest frequency point. The corresponding relationship represents whether the historical K song audio of the target user meets the original song audio.

[0064] Additionally, it should be noted that, since the audio frequency point corresponding to the first audio data of the target user direction before adjustment and the audio frequency point corresponding to the second audio data of the non-target user direction are both preset default frequency points, and the second preset audio frequency point is less than the first preset audio frequency point, there are two cases for amplifying the target user direction of the human voice ratio, that is, when the user's singing level sings the song without difficulty, the song audio frequency point can be used as the first preset audio frequency point, and when the user's singing level sings the song with difficulty, the user audio frequency point can be used as the first preset audio frequency point. Because the reference frequency point is different, the first preset audio frequency point is also different, so the user audio frequency point is used to represent the audio frequency point of the song corresponding to the song audio data sung by the target user, wherein the first preset audio frequency point and the second first preset audio frequency point can be one or more. When the first preset audio frequency point is multiple, it means that the smart display device locally enhances the song segment corresponding to the song audio data.

[0065] As an example, steps C10 to C30 include: collecting song audio data and historical karaoke audio data of the target user in the karaoke mode through a microphone array of the smart display device; obtaining a corresponding relationship between a user audio frequency point corresponding to the historical karaoke audio data and a song audio frequency point corresponding to the song audio data; and determining the first preset audio frequency point according to the corresponding relationship and a song segment corresponding to the song audio data.

[0066] The step of determining the first preset audio frequency point according to the corresponding relationship and the song segment corresponding to the song audio data includes:

[0067] Step D10, detecting a song segment type of the song segment;

[0068] Step D20, if the song segment is a first type of song segment, then the song audio frequency point is used as the first preset audio frequency point;

[0069] Step D30, if the song segment is a second type of song segment, then when the corresponding relationship is a first corresponding relationship, the song audio frequency point is used as the first preset audio frequency point;

[0070] Step D40, when the corresponding relationship is a second corresponding relationship, the user audio frequency point is used as the first preset audio frequency point.

[0071] In this embodiment, it should be noted that the first type of song fragment is used to represent that the song fragment has no singing difficulty, the second type of song fragment is used to represent that the song has a certain singing difficulty, the first correspondence is used to represent that the target user's historical karaoke audio does not meet the original audio of the song, and the second correspondence is used to represent that the target user's historical karaoke audio meets the original audio of the song. For example, in one implementable method, assuming that the current song fragment interval of the original audio of the song has a high-pitched audio frequency point and a low-pitched audio frequency point, then the lowest audio frequency point of the target user's historical karaoke audio is lower than the low-pitched audio frequency point of the song, and the highest high-frequency point of the historical karaoke audio is higher than the high-pitched audio frequency point of the song, then the correspondence is determined to be the first correspondence.

[0072] As an example, steps D10 to D20 include: detecting the song segment type of the song segment; if the song segment is a song segment without singing difficulty, then the song audio frequency point is used as the first preset audio frequency point; if the song segment is a song segment with singing difficulty, then when the correspondence is a first correspondence, the song audio frequency point is used as the first preset audio frequency point, and when the correspondence is a second correspondence, the user audio frequency point is used as the first preset audio frequency point.

[0073] The step of determining the first audio data of the target user's direction in the multi-person voice audio data includes:

[0074] Step E10: Obtain the user distance of the target user and the preset number of audio signal transmission and reception times corresponding to the multi-person audio data;

[0075] Step E20: Calculate the target signal transmission and reception time of the target user based on the user distance;

[0076] Step E30: Determine the user's karaoke audio based on the time difference between the target signal transmission and reception time and the transmission and reception time of each audio signal.

[0077] In the embodiment, it is to be explained that, since the proportion of human voice in the direction of the target user is to be enhanced, the target user needs to be positioned, which can be realized by the millimeter wave radar installed in the smart display device. The user distance is used to represent the relative distance between the target user and the smart display device, and the target signal transmission time is used to represent the time when the millimeter wave signal transmitted by the millimeter wave radar is reflected and received by the target user. When the human voice audio data is the multi-person audio data, that is, the microphone array collects the human voice audio data of multiple users respectively, since the direction positions of different users and the smart display device are different, and the time when the millimeter wave signal transmitted by the millimeter wave radar is reflected and received by the user is different, the audio signal transmission time is used to represent the time when the millimeter wave signal transmitted by the millimeter wave radar is reflected and received by the user.

[0078] As an example, steps E10 to E30 include: acquiring the user distance of the target user, and acquiring a preset number of audio signal transmission times corresponding to the multi-person audio data; calculating the target signal transmission time of the target user according to the user distance; and taking the human voice audio data corresponding to the audio signal transmission time with the minimum time difference between the target signal transmission time and each audio signal transmission time as the karaoke audio. Since the target signal transmission time is the theoretical calculation time of the target user reflecting the millimeter wave signal transmitted by the millimeter wave radar to the millimeter wave radar, when there are multiple users emitting human voice audio data, the target user can be accurately determined according to the difference between the actual calculation time and the theoretical calculation time of the target user reflecting the millimeter wave signal transmitted by the millimeter wave radar to the millimeter wave radar, and the audio data in the direction of the target user, thereby laying a foundation for adjusting the karaoke audio data.

[0079] The embodiment of the application provides a karaoke method, which is applied to a smart display device, that is, obtaining a lip region image of a target user in a karaoke mode; determining whether the target user is in a singing state by image recognition on the lip region image, that is, realizing the purpose of determining whether the user is in the karaoke by image recognition on the lip region image of the target user; and then if yes, collecting karaoke audio data corresponding to the target user in the karaoke mode, wherein the karaoke audio data includes first audio data in the direction of the target user and second audio data in the direction of non-target user; then adjusting the first audio data by a first preset audio frequency point and adjusting the second audio data by a second preset audio frequency point to obtain to-be-played audio data, wherein the first preset audio frequency point is greater than the second preset audio frequency point; and then playing the to-be-played audio data, so that the target user and the smart display device interact in the karaoke according to the to-be-played audio data. Since the audio frequency point corresponding to the audio data in the direction of the target user and the audio frequency point corresponding to the audio data in the direction of non-target user are consistent when the target user is in the karaoke mode, the first audio data in the direction of the target user is enhanced by the first preset audio frequency point, and the second audio data in the direction of non-target user is weakened by the second preset audio frequency point, so that the purpose of enhancing the audio data in the target direction is realized, that is, the karaoke voice ratio of the target user is directly amplified, and the expected karaoke effect is achieved, that is, even if the user does not purchase a karaoke device, the karaoke voice ratio can be amplified when the user is in the karaoke, and when the user is in the karaoke mode and is singing, the karaoke voice ratio needs to be amplified indirectly by a microphone or other karaoke external device, so that the technical defect that the karaoke effect is poor due to the consideration of the karaoke cost without purchasing the karaoke device is overcome, and the purpose of considering the karaoke effect and the karaoke cost is realized.

[0080] Embodiment two

[0081] Further, with reference to Figure 3 In another embodiment of the application, the same or similar contents as the above embodiment one can be referred to the above introduction, and the subsequent will not be described. On this basis, the step of judging whether the target user is in a singing state according to the first image feature and the second image feature comprises:

[0082] Step F10, calculating the feature similarity between the first image feature and the second image feature, and detecting whether the feature similarity is greater than a preset feature similarity threshold;

[0083] Step F20, if greater, determining that the target user is not in a singing state;

[0084] Step F30, if not greater, determining that the target user is in a singing state.

[0085] In the embodiment, it is to be explained that, since there is a certain limitation in the similarity comparison between the two images by naked eyes, the consistency of the two images can be determined by comparing the feature similarity, the feature similarity is used to represent the similarity between the first lip key frame image and the second lip key frame image, and can be specifically represented by the distance between the features, wherein the distance between the features can be Manhattan distance or Minkowski distance, etc., and the preset feature similarity threshold is a pre-set similarity critical value for representing the consistency between the first lip key frame image and the second lip key frame image.

[0086] As an example, steps F10 to F40 include: calculating the feature similarity between the first image feature and the second image feature, and comparing whether the feature similarity is greater than a preset feature similarity threshold; if the feature similarity is greater than the preset feature similarity threshold, it is determined that the first lip key frame image and the second lip key frame image are consistent, and it is further determined that the target user is not in a singing state; if the feature similarity is not greater than the preset feature similarity threshold, it is determined that the first lip key frame image and the second lip key frame image are inconsistent, and it is further determined that the target user is in a singing state.

[0087] The embodiment of the present application provides a singing state judgment method, that is, calculating the feature similarity between the first image feature and the second image feature, and detecting whether the feature similarity is greater than a preset feature similarity threshold; if it is greater, it is determined that the target user is not in a singing state; if it is not greater, it is determined that the target user is in a singing state. Compared with the judgment method of comparing the similarity between the first lip key frame image and the second lip key frame image by naked eyes, and then determining whether the target user is in a singing state, the embodiment of the present application judges whether the target user is in a singing state based on the feature similarity of the two images. Since the feature similarity can accurately reflect the similarity between the two images, even if the target user sings karaoke by slightly changing the lips, it can be accurately determined, so the accuracy of the singing state judgment of the target user is improved.

[0088] Embodiment three

[0089] The embodiment of the present application also provides a karaoke device applied to an intelligent display device, referring to Figure 4 , the karaoke device comprises:

[0090] The acquisition module 101 is configured to acquire the lip region image of the target user in the karaoke mode.

[0091] The determining module 102 is configured to determine whether the target user is in a singing state by performing image recognition on the lip region image.

[0092] The collecting module 103 is configured to collect K-song audio data corresponding to the target user in the K-song mode if the target user is in the singing state, wherein the K-song audio data includes first audio data in the direction of the target user and second audio data in the direction of non-target users.

[0093] The adjusting module 104 is configured to adjust the first audio data by a first preset audio frequency point and adjust the second audio data by a second preset audio frequency point to obtain to-be-played audio data, wherein the first preset audio frequency point is greater than the second preset audio frequency point.

[0094] The playing module 105 is configured to play the to-be-played audio data for the target user and the intelligent display device to perform K-song interaction according to the to-be-played audio data.

[0095] Optionally, the lip region image includes a first lip key frame image and a second lip key frame image, and the determining module 102 is further configured to:

[0096] extract image features from the first lip key frame image and the second lip key frame image based on a preset image feature extraction model to obtain first image features and second image features;

[0097] determine whether the target user is in a singing state according to the first image features and the second image features.

[0098] Optionally, the interface description information includes identification description information and video description information, and the determining module 102 is further configured to:

[0099] calculate a feature similarity between the first image features and the second image features, and detect whether the feature similarity is greater than a preset feature similarity threshold;

[0100] if greater, it is determined that the target user is not in a singing state;

[0101] if not greater, it is determined that the target user is in a singing state.

[0102] Optionally, the collecting module 103 is further configured to:

[0103] collect song external audio data in the K-song mode, wherein the song external audio data includes human voice audio data and interference audio data;

[0104] If the human voice audio data is single-person voice audio data, the single-person voice audio data is taken as first audio data of the target user direction, and the interference audio data is taken as second audio data of the non-user direction;

[0105] If the human voice audio data is multi-person voice audio data, first audio data of the target user direction is determined in the multi-person voice audio data;

[0106] Second audio data of the non-user direction is determined according to the first audio data and the interference audio data.

[0107] Optionally, the K-song device is further used for:

[0108] In the K-song mode, song audio data and historical K-song audio data of the target user are collected;

[0109] A corresponding relationship between user audio frequency points corresponding to the historical K-song audio data and song audio frequency points corresponding to the song audio data is obtained;

[0110] The first preset audio frequency point is determined according to the corresponding relationship and a song segment corresponding to the song audio data.

[0111] Optionally, the K-song device is further used for:

[0112] The K-song device is further used for.

[0113] Optionally, the collection module 103 is further used for:

[0114] A user distance of the target user is obtained, and a preset number of audio signal transmission / reception times corresponding to the multi-person voice audio data are obtained;

[0115] A target signal transmission / reception time of the target user is calculated according to the user distance;

[0116] The user K-song audio is determined according to the target signal transmission / reception time and a time difference between each audio signal transmission / reception time.

[0117] The K-song device provided by the present application adopts the K-song method in the above embodiments, and solves the technical problem that home K-song is difficult to balance K-song effect and K-song cost. Compared with the prior art, the K-song device provided by the present application has the same beneficial effects as the K-song method provided by the above embodiments, and other technical features in the K-song device are the same as the features disclosed in the above embodiments, which will not be repeated here.

[0118] Embodiment Four

[0119] The electronic device provided by the embodiment of the present application comprises: at least one processor; and a memory connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the K-song method in the above-mentioned embodiment I.

[0120] Reference will now be made to the drawings Figure 5 , which show structural schematic diagrams of an electronic device suitable for implementing the embodiments of the present disclosure. The electronic device in the embodiments of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), and the like, as well as fixed terminals such as digital TVs, desktop computers, and the like. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.

[0121] As shown in Figure 5 , the electronic device can include a processing device 1001 (such as a central processor, a graphics processor, etc.) that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1002 or loaded from a storage device 1003 into a random access memory (RAM) 1004. Various programs and data required for the operation of the electronic device are also stored in the RAM 1004. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus.

[0122] Generally, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 1009. The communication devices can allow the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although the electronic device is shown with various systems, it should be understood that all of the systems shown are not required to be implemented or possessed. More or fewer systems can be alternatively implemented or possessed.

[0123] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication device 1009, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0124] The electronic device provided by the present application adopts the K song method in the above-mentioned embodiments, and solves the technical problem that it is difficult to balance the K song effect and K song cost in home K song. Compared with the prior art, the electronic device provided by the embodiments of the present application has the same beneficial effects as the K song method provided by the above-mentioned embodiments, and other technical features in the electronic device are the same as the features disclosed in the above-mentioned embodiments, which will not be repeated here.

[0125] It should be understood that parts of the present disclosure can be realized by hardware, software, firmware or a combination thereof. In the description of the above-mentioned embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0126] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0127] Embodiment five

[0128] The present embodiment provides a computer readable storage medium having stored thereon computer readable program instructions for performing the K song method in the above-mentioned embodiments.

[0129] The computer readable storage medium provided by the embodiment of the present application may, for example, be a U disk, but is not limited to an electric, magnetic, optical, electromagnetic, infrared, or semiconductor system, system, or device, or any combination thereof. More specific examples of the computer readable storage medium may include, but are not limited to, an electric connection with one or more conductive wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the embodiment, the computer readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system, or device. The program code contained on the computer readable storage medium can be transmitted by any suitable medium, including but not limited to an electric wire, an optical cable, an RF (radio frequency), and the like, or any suitable combination thereof.

[0130] The computer readable storage medium described above may be contained in an electronic device, or may exist separately without being assembled into an electronic device.

[0131] The computer readable storage medium described above carries one or more programs, which, when executed by an electronic device, cause the electronic device to: acquire a lip region image of a target user in a K song mode; determine whether the target user is in a singing state by performing image recognition on the lip region image; if so, collect K song audio data corresponding to the target user in the K song mode, wherein the K song audio data includes first audio data in the direction of the target user and second audio data in the direction of a non-target user; adjust the first audio data by a first preset audio frequency point and adjust the second audio data by a second preset audio frequency point to obtain to-be-played audio data, wherein the first preset audio frequency point is greater than the second preset audio frequency point; and play the to-be-played audio data for K song interaction between the target user and the smart display device based on the to-be-played audio data.

[0132] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0133] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0134] The modules involved in the embodiments of the present disclosure can be implemented in the manner of software or hardware. In some cases, the name of the module does not constitute a limitation on the module itself.

[0135] The computer readable storage medium provided by the present application stores computer readable program instructions for executing the K singing method, and solves the technical problem that the home K singing is difficult to balance the K singing effect and the K singing cost. Compared with the prior art, the computer readable storage medium provided by the embodiment of the present application has the same beneficial effects as the K singing method provided by the above-mentioned embodiment, and will not be described here.

[0136] Embodiment six

[0137] The present application also provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of the K singing method as described above.

[0138] The computer program product provided in the application solves the technical problem that home KTV is difficult to balance KTV effect and KTV cost. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiment of the application are the same as those of the KTV method provided in the above-mentioned embodiment, and are not described herein.

[0139] The above is only the preferred embodiment of the application, and does not limit the patent scope of the application, and any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent processing scope of the application.

Claims

1. A karaoke method, characterized in that, The karaoke method, applied to smart display devices, includes: In karaoke mode, obtain an image of the target user's lip area; By performing image recognition on the lip region image, it can be determined whether the target user is in a singing state; If so, then in the karaoke mode, collect the karaoke audio data corresponding to the target user, wherein the karaoke audio data includes first audio data in the direction of the target user and second audio data in the direction of the non-target user; The method involves adjusting the first audio data using a first preset audio frequency point and adjusting the second audio data using a second preset audio frequency point to obtain audio data to be played, wherein the first preset audio frequency point is greater than the second preset audio frequency point. Before the step of adjusting the first audio data using the first preset audio frequency point and adjusting the second audio data using the second preset audio frequency point to obtain the audio data to be played, the karaoke method includes: In the karaoke mode, song audio data and the target user's historical karaoke audio data are collected; the correspondence between the user's audio frequency point corresponding to the historical karaoke audio data and the song's audio frequency point corresponding to the song audio data is obtained; based on the correspondence and the song segment corresponding to the song audio data, the first preset audio frequency point is determined. When singing a song is not difficult, the song's audio frequency point is used as the first preset audio frequency point; when singing a song is difficult, the user's audio frequency point is used as the first preset audio frequency point. The audio data to be played is played so that the target user and the smart display device can interact by singing karaoke based on the audio data to be played.

2. The karaoke method as described in claim 1, characterized in that, The lip region image includes a first lip keyframe image and a second lip keyframe image. The step of determining whether the target user is singing by performing image recognition on the lip region image includes: Based on a preset image feature extraction model, image features are extracted from the first lip keyframe image and the second lip keyframe image respectively to obtain the first image features and the second image features; Based on the first image features and the second image features, it is determined whether the target user is in a singing state.

3. The karaoke method as described in claim 2, characterized in that, The step of determining whether the target user is singing based on the first image feature and the second image feature includes: Calculate the feature similarity between the first image feature and the second image feature, and detect whether the feature similarity is greater than a preset feature similarity threshold; If the value is greater than the target user's, then the target user is determined not to be in a singing state. If the value is not greater than the value, then the target user is determined to be in a singing state.

4. The karaoke method as described in claim 1, characterized in that, The step of collecting the karaoke audio data corresponding to the target user in the karaoke mode includes: In the karaoke mode, external audio data of the song is collected, wherein the external audio data of the song includes human voice audio data and interference audio data; If the human voice audio data is single-voice audio data, then the single-voice audio data is used as the first audio data in the direction of the target user, and the interference audio data is used as the second audio data in the direction of the non-target user. If the human voice audio data is multi-voice audio data, then determine the first audio data of the target user direction from the multi-voice audio data; Based on the first audio data and the interference audio data, the second audio data for the non-target user direction is determined.

5. The karaoke method as described in claim 1, characterized in that, The step of determining the first preset audio frequency point based on the correspondence and the song segment corresponding to the song audio data includes: Detect the song segment type of the song segment; If the song fragment is a first type of song fragment, then the song audio frequency point is taken as the first preset audio frequency point; If the song fragment is a second type of song fragment, then when the correspondence is a first correspondence, the song audio frequency point is taken as the first preset audio frequency point; When the correspondence is the second correspondence, the user audio frequency point is used as the first preset audio frequency point.

6. The karaoke method as described in claim 4, characterized in that, The step of determining the first audio data of the target user's direction in the multi-person voice audio data includes: Obtain the user distance of the target user, and obtain the audio signal transmission and reception time of a preset number of audio signals corresponding to the multi-person audio data; Based on the user distance, calculate the target signal transmission and reception time for the target user; The user's karaoke audio is determined based on the time difference between the target signal transmission and reception time and the transmission and reception time of each audio signal.

7. A karaoke device, characterized in that, The karaoke device, applied to smart display devices, includes: The acquisition module is used to acquire images of the lip region of the target user in karaoke mode; The determination module is used to determine whether the target user is singing by performing image recognition on the lip region image; The acquisition module is used to acquire, if yes, the karaoke audio data corresponding to the target user in the karaoke mode, wherein the karaoke audio data includes first audio data in the direction of the target user and second audio data in the direction of non-target user; An adjustment module is used to adjust the first audio data using a first preset audio frequency point and to adjust the second audio data using a second preset audio frequency point to obtain audio data to be played, wherein the first preset audio frequency point is greater than the second preset audio frequency point. Before the step of adjusting the first audio data using the first preset audio frequency point and adjusting the second audio data using the second preset audio frequency point to obtain the audio data to be played, the karaoke method includes: In the karaoke mode, song audio data and the target user's historical karaoke audio data are collected; the correspondence between the user's audio frequency point corresponding to the historical karaoke audio data and the song's audio frequency point corresponding to the song audio data is obtained; based on the correspondence and the song segment corresponding to the song audio data, the first preset audio frequency point is determined. When singing a song is not difficult, the song's audio frequency point is used as the first preset audio frequency point; when singing a song is difficult, the user's audio frequency point is used as the first preset audio frequency point. The playback module is used to play the audio data to be played, so that the target user and the smart display device can perform karaoke interaction based on the audio data to be played.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the steps of the karaoke method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program for implementing the karaoke method, which is executed by a processor to implement the steps of the karaoke method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for displaying sound correcting state

    CN108492807A

  • Method and device for adjusting sound effect of vehicle-mounted sound system

    CN114734942A