Sound production position determination method and apparatus, and computer device

By acquiring and adjusting the sound information of video frames, including object identification, confidence level, and location information, the problem of sound location delay caused by neural network recognition time is solved, and more accurate sound location matching is achieved.

WO2025251770A1PCT designated stage Publication Date: 2025-12-11HUIZHOU VISION NEW TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/087329
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-04
Filing Date
2025-04-03
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

In the current process of sound objectification, it takes time for the neural network to identify whether an object in the current scene is making a sound, which leads to a mismatch between the sound location and the scene, resulting in a delay in the effect.

Method used

By acquiring the sound information of the current video frame, including object identification information, sound confidence information, and object location information, the number of objects is determined. Based on the object location information and sound confidence information, the initial sound position is adjusted to obtain the target sound position, reducing the delay caused by recognition time.

Benefits of technology

It improves the matching degree between the sound position and the image, alleviates the effect delay caused by the sound recognition time, and achieves more accurate sound position determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025087329_11122025_PF_FP_ABST
    Figure CN2025087329_11122025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a sound production position determination method and apparatus, and a computer device. The method comprises: on the basis of object identification information, determining the number of objects in a current video frame; if the number of objects is greater than or equal to a first number threshold and less than a second number threshold, determining initial sound production position information of the current video frame on the basis of object position information; and on the basis of sound production confidence information and sound production flag bit information, adjusting the initial sound production position information to obtain target sound production position information.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for determining sound emitting position and computer device

[0001] The present application claims priority to the Chinese patent application No. 202410718254.0, filed on June 4, 2024, and entitled "Method and device for determining sound emitting position, computer device and storage medium", the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the technical field of computer, in particular to a method and device for determining sound emitting position and a computer device. BACKGROUND

[0003] Sound objectification refers to analyzing the motion coordinates and trajectories of different sound emitting objects from a picture by a specific method, and then matching the labeled sound to the corresponding picture object through sound separation and sound labeling, so as to convert the stereophonic sound into a plurality of audio metadata with motion trajectories and audio data. TECHNICAL PROBLEM

[0004] In the existing sound objectification process, a neural network is usually used to identify whether the current picture object emits sound, and the sound emitting position of the picture is adjusted after the current picture object is identified to emit sound. Since the neural network needs a period of time to identify whether the current picture object emits sound, the sound emitting position and the picture do not match in this period of time. TECHNICAL SOLUTION

[0005] The embodiments of the present application provide a method and device for determining sound emitting position and a computer device, which can obtain sound emitting position information that is more matched with the picture, and can alleviate the effect delay caused by sound emitting identification time.

[0006] In a first aspect, the present application provides a method for determining sound emitting position, comprising:

[0007] obtaining sound emitting information of a current video frame; the sound emitting information comprises object identification information, sound emitting confidence information, sound emitting flag information and object position information of an object included in the current video frame;

[0008] determining the number of objects in the current video frame based on the object identification information;

[0009] if the number of objects is greater than or equal to a first number threshold and less than a second number threshold, determining initial sound emitting position information of the current video frame based on the object position information;

[0010] adjusting the initial sound emitting position information based on the sound emitting confidence information and the sound emitting flag information to obtain target sound emitting position information of the current video frame.

[0011] In some embodiments of the present application, the initial speech position information is adjusted based on the speech confidence information and the speech flag information to obtain target speech position information of the current video frame, including:

[0012] Speech recognition is performed based on the speech confidence information and the speech flag information to obtain speech state information of the target object; the target object is an object with the highest priority among objects included in the current video frame;

[0013] If the speech state information is a first speech state and the target object is a person, the object position information of the target object is compared with a first position threshold and a second position threshold; the first speech state indicates that the target object is speaking;

[0014] If the object position information of the target object is greater than the first position threshold and less than the second position threshold, the initial speech position information is adjusted based on a face deflection angle of the target object to obtain candidate speech position information;

[0015] Based on the candidate speech position information, target speech position information of the current video frame is determined.

[0016] In some embodiments of the present application, the initial speech position information is adjusted based on a face deflection angle of the target object to obtain candidate speech position information, including:

[0017] The face deflection angle of the target object is compared with a first angle threshold, a second angle threshold, a third angle threshold, and a fourth angle threshold;

[0018] If the face deflection angle is greater than or equal to the first angle threshold and less than the second angle threshold, first position information is determined based on the face deflection angle and the first angle threshold;

[0019] The first position information is added to preset position information to obtain second position information;

[0020] The initial speech position information is subtracted from the second position information to obtain the candidate speech position information.

[0021] In some embodiments of the present application, after the face deflection angle of the target object is compared with the first angle threshold, the second angle threshold, the third angle threshold, and the fourth angle threshold, the method further includes:

[0022] If the face deflection angle is greater than the third angle threshold and less than or equal to the fourth angle threshold, the first position information is determined based on the face deflection angle and the first angle threshold;

[0023] The first position information is subtracted from the preset position information to obtain third position information;

[0024] Subtracting the initial sound position information from the third position information, candidate sound position information is obtained.

[0025] In some embodiments of the present application, sound recognition is performed based on the sound confidence information and the sound flag bit information, to obtain sound state information of the target object, including:

[0026] Obtaining scene switching identification information, a preset confidence threshold, a first reference value and a second reference value.

[0027] If the sound confidence information is less than the preset confidence threshold, the sound flag bit information is not the first reference value, and the scene switching identification information is the second reference value, it is determined that the sound state information of the target object is a second sound state; the second sound state represents that the target object does not make a sound.

[0028] If the sound confidence information is greater than or equal to the preset confidence threshold, and / or the sound flag bit information is the first reference value, and / or the scene switching identification information is not the second reference value, it is determined that the sound state information of the target object is a first sound state.

[0029] In some embodiments of the present application, after obtaining the sound state information of the target object based on the sound confidence information and the sound flag bit information, the method further includes:

[0030] If the sound state information is the second sound state, updating a post-sound judgment queue value.

[0031] Determining whether the updated post-sound judgment queue value is greater than a queue upper limit value.

[0032] If the updated post-sound judgment queue value is less than or equal to the queue upper limit value, obtaining target sound position information of a historical video frame, and determining target sound position information of a current video frame based on the target sound position information of the historical video frame and the initial sound position information.

[0033] If the updated post-sound judgment queue value is greater than the queue upper limit value, determining a preset sound position information as the initial sound position information, and determining target sound position information of a current video frame based on the target sound position information of the historical video frame and the initial sound position information.

[0034] In some embodiments of the present application, based on the object identification information, the number of objects in the current video frame is determined, including:

[0035] Based on the object identification information and preset object priority information, a target object and an identification number of the target object are determined; the target object is an object with the highest priority among the objects included in the current video frame.

[0036] Based on the identification number of the target object, the number of target objects is determined.

[0037] The number of target objects is determined as the number of objects in the current video frame.

[0038] In some embodiments of the present application, after determining the number of objects in the current video frame based on the object identification information, the method further comprises:

[0039] If the number of objects is less than the first number threshold or greater than or equal to the second number threshold, the preset sound emission position information is determined as the initial sound emission position information of the current video frame.

[0040] The target sound emission position information of the historical video frame is obtained, and the target sound emission position information of the current video frame is determined based on the target sound emission position information of the historical video frame and the initial sound emission position information.

[0041] In some embodiments of the present application, the sound emission information of the current video frame is obtained, comprising:

[0042] The image recognition model is used to perform image recognition on the current video frame to obtain object information of the objects contained in the current video frame.

[0043] The object information is converted to obtain the sound emission information of the current video frame.

[0044] In some embodiments of the present application, the initial sound emission position information of the current video frame is determined based on the object position information, comprising:

[0045] The object position information of the target object is determined from the object position information.

[0046] The initial sound emission position information of the current video frame is determined based on the object position information of the target object.

[0047] In some embodiments of the present application, the sound emission confidence information in the sound emission information is obtained by the following steps:

[0048] The time sequence information of the landmark of the object is determined based on the landmark information of the object.

[0049] The class variance data is determined based on the time sequence information.

[0050] The sound emission confidence information is determined based on the class variance data.

[0051] In some embodiments of the present application, the sound emission flag information in the sound emission information is obtained by the following steps:

[0052] The sound emission confidence information is compared with a preset first threshold.

[0053] If the sound emission confidence information is greater than the preset first threshold, the sound emission flag information is determined as a first reference value.

[0054] If the voice generation confidence information is less than or equal to a preset first threshold value, the voice generation flag information is determined as a third reference value.

[0055] In some embodiments of the present application, after comparing the face deflection angle of the target object with the first angle threshold value, the second angle threshold value, the third angle threshold value and the fourth angle threshold value, the method further comprises:

[0056] If the face deflection angle is greater than the fourth angle threshold value and less than the first angle threshold value, the initial voice generation position information is determined as the candidate voice generation position information.

[0057] In some embodiments of the present application, if the updated post-voice generation judgment queue value is greater than the upper queue limit value, the method further comprises:

[0058] The post-voice generation judgment queue value is set to the sum of the upper queue limit and a preset fourth threshold value.

[0059] In some embodiments of the present application, if the voice generation state information is the first voice generation state, the method further comprises:

[0060] The post-voice generation judgment queue value is set to the fourth reference value, and the scene switching identification information is set to the fourth reference value.

[0061] In some embodiments of the present application, if the object quantity is greater than or equal to the first quantity threshold value and less than the second quantity threshold value, the method further comprises:

[0062] Obtaining the difference between the initial voice generation position information and the target voice generation position information of the previous video frame;

[0063] If the difference is greater than a preset fifth threshold value, the scene switching identification information is set to a second reference value.

[0064] In some embodiments of the present application, if the face deflection angle is greater than or equal to the first angle threshold value and less than the second angle threshold value, the calculation process of the candidate voice generation position information is: pos2 = pos1 + (A - B) * angle wherein pos2 represents the candidate voice generation position information, pos1 represents the initial voice generation position information, A represents the first angle threshold value, B represents the preset position information, and angle represents the face deflection angle.

[0065] In some embodiments of the present application, if the face deflection angle is greater than the third angle threshold value and less than or equal to the fourth angle threshold value, the calculation process of the candidate voice generation position information is: pos2 = pos1 + (A - B) * angle wherein pos2 represents the candidate voice generation position information, pos1 represents the initial voice generation position information, A represents the first angle threshold value, B represents the preset position information, and angle represents the face deflection angle.

[0066] In a second aspect, the embodiments of the present application further provide a sound source position determination apparatus, comprising:

[0067] an information obtaining module, configured to obtain sound source information of a current video frame; the sound source information comprises object identification information, sound source confidence information, sound source flag information and object position information of an object included in the current video frame;

[0068] a quantity determining module, configured to determine the number of objects in the current video frame based on the object identification information;

[0069] a position determining module, configured to determine initial sound source position information of the current video frame based on the object position information if the number of objects is greater than or equal to a first number threshold and less than a second number threshold;

[0070] a position adjusting module, configured to adjust the initial sound source position information based on the sound source confidence information and the sound source flag information to obtain target sound source position information of the current video frame.

[0071] In a third aspect, the present application further provides a computer device, comprising:

[0072] one or more processors;

[0073] a memory; and

[0074] one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the processor to implement the sound source position determination method of any one of the first aspect.

[0075] In a fourth aspect, the embodiments of the present application provide a computer readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to execute the steps in the sound source position determination method of any one of the first aspect. Advantages

[0076] The number of objects in the current video frame is determined based on the object identification information, the initial sound source position information of the current video frame is determined based on the object position information if the number of objects is greater than or equal to the first number threshold and less than the second number threshold, the initial sound source position information is adjusted based on the sound source confidence information and the sound source flag information to obtain the target sound source position information of the current video frame, and when it is not identified whether the object in the picture makes a sound, the sound source position is determined based on the object position information first, and then the sound recognition and the initial sound source position adjustment are performed based on the sound source confidence information and the sound source flag information, so that not only the sound source position information that is more matched with the picture can be obtained, but also the effect delay caused by the sound recognition time can be relieved. BRIEF DESCRIPTION OF DRAWINGS

[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without any creative effort.

[0078] Fig. 1 is a flow diagram of a sound emitting position determination system according to an embodiment of the present application;

[0079] Fig. 2 is a flow diagram of a sound emitting position determination method according to an embodiment of the present application;

[0080] Fig. 3 is a structural diagram of screen division according to an embodiment of the present application;

[0081] Fig. 4 is a specific flow diagram of obtaining sound emitting information according to an embodiment of the present application;

[0082] Fig. 5 is a specific flow diagram of determining the number of objects according to an embodiment of the present application;

[0083] Fig. 6 is a specific flow diagram of adjusting initial sound emitting position information according to an embodiment of the present application;

[0084] Fig. 7 is a principle block diagram of a sound emitting position determination apparatus according to an embodiment of the present application;

[0085] Fig. 8 is a structural diagram of one embodiment of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION

[0086] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort fall within the scope of the present application.

[0087] In the description of the present application, it needs to be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second", "third" and the like are only for description purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined with "first", "second", "third" and the like can explicitly or implicitly include one or more of the features.

[0088] In the present application, the word "exemplary" is used to mean "serving as an example, instance, or illustration." Any implementation described as "exemplary" in the present application is not necessarily to be construed as preferred or advantageous over other implementations. The following description is presented to enable any person skilled in the art to make and use the present application. In the following description, for purposes of explanation, specific details are set forth. It is apparent to those skilled in the art that the present application can be practiced without using these specific details. In other instances, well-known structures and processes are not described in detail in order to avoid obscuring the description of the present application. Thus, the present application is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features presented herein.

[0089] It should be noted that the method of the present application is executed in a computer device, and the processing objects of each computer device exist in the form of data or information, such as time, which is essentially time information. It can be understood that in subsequent embodiments, if the size, quantity, position and the like are mentioned, they are all corresponding data, so that the computer device can process them, and specific details are not described here.

[0090] The present application provides a sound position determination method and device, and a computer device, which are described in detail below.

[0091] Please refer to FIG. 1, which is a scene schematic diagram of a sound position determination system provided by an embodiment of the present application. The sound position determination system can include a computer device 100, and the computer device 100 is integrated with a sound position determination device, such as the computer device in FIG. 1.

[0092] The computer device 100 in the embodiments of the present application is mainly used to obtain sound emission information of a current video frame; the sound emission information comprises object identification information, sound emission confidence information, sound emission flag bit information and object position information of an object contained in the current video frame; based on the object identification information, the number of objects in the current video frame is determined; if the number of objects is greater than or equal to a first number threshold and less than a second number threshold, initial sound emission position information of the current video frame is determined based on the object position information; the initial sound emission position information is adjusted based on the sound emission confidence information and the sound emission flag bit information to obtain target sound emission position information of the current video frame, which can improve the picture contrast of the display terminal from the signal dimension and the backlight dimension.

[0093] In the embodiments of the present application, the computer device 100 can be a stand-alone server, or a server network or a server cluster composed of servers. For example, the computer device 100 described in the embodiments of the present application includes but is not limited to a computer, a network host, a single network server, a plurality of network server sets or a cloud server composed of a plurality of servers. The cloud server is composed of a large number of computers or network servers based on cloud computing.

[0094] It can be understood that the computer device 100 used in the embodiments of the present application can be a device that includes receiving and transmitting hardware, i.e., a device with receiving and transmitting hardware capable of performing bidirectional communication on a bidirectional communication link. Such a device can include a cellular or other communication device with a single-line display or a multi-line display or a cellular or other communication device without a multi-line display. The specific computer device 100 can be a desktop terminal or a mobile terminal, and the computer device 100 can also be one of a smart television, a mobile phone, a tablet computer, a notebook computer, etc.

[0095] Those skilled in the art can understand that the application environment shown in FIG. 1 is only one application scenario of the present application scheme, and does not constitute a limitation on the application scenarios of the present application scheme. Other application environments can include more or fewer computer devices than those shown in FIG. 1. For example, only one computer device is shown in FIG. 1. It can be understood that the sound emission position determination system can also include one or more other services, which are not limited here.

[0096] In addition, as shown in FIG. 1, the sound emission position determination system can also include a memory 200 for storing data, such as sound emission information, e.g., object identification information, sound emission confidence information, sound emission flag bit information, etc., and sound emission position information, e.g., initial sound emission position information, target sound emission position information, preset sound emission position information, etc.

[0097] It should be noted that the scene diagram of the sound position determination system shown in FIG. 1 is only an example, and the sound position determination system and the scene described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of the sound position determination system and the appearance of new business scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0098] First, a sound position determination method is provided in the embodiments of the present application. The execution subject of the sound position determination method is a sound position determination device. The sound position determination device is applied to a computer device. The sound position determination method comprises: obtaining sound information of a current video frame; the sound information comprises object identification information, sound confidence information, sound flag bit information and object position information of an object included in the current video frame; determining the number of objects in the current video frame based on the object identification information; if the number of objects is greater than or equal to a first number threshold and less than a second number threshold, determining initial sound position information of the current video frame based on the object position information; and adjusting the initial sound position information based on the sound confidence information and the sound flag bit information to obtain target sound position information of the current video frame.

[0099] As shown in FIG. 2, it is a flowchart of an embodiment of the sound position determination method in the embodiments of the present application. The sound position determination method can comprise the following steps S201-S204, which are as follows.

[0100] Step S201, obtaining sound information of a current video frame; the sound information comprises object identification information, sound confidence information, sound flag bit information and object position information of an object included in the current video frame.

[0101] Specifically, the current video frame is a video frame that needs to be determined for sound position, the sound information is information related to sound in the current video frame, the object included in the current video frame is a target object in the current video frame, and the object includes but is not limited to a person, a machine, a musical instrument, an animal, etc.

[0102] The sound information comprises object identification information, sound confidence information, sound flag bit information and object position information of the object included in the current video frame. The object identification information represents the object category and the object number of the object included in the current video frame. For example, object identification 00:1 represents a person, object identification 00:2 represents a machine, object identification 00:3 represents a musical instrument, and object identification 00:4 represents an animal. If the object identification information of the object included in the current video frame is 00:1, 00:1, 00:2 and 00:4, it means that the current video frame includes two persons, one machine and one animal.

[0103] Further, the voice generation confidence information represents a possibility of voice generation of the object, and the greater the voice generation confidence information is, the higher the possibility of voice generation of the object is, and the smaller the voice generation confidence information is, the lower the possibility of voice generation of the object is. The voice generation flag information is determined based on the voice generation confidence information. If the voice generation confidence information is greater than a preset first threshold, the voice generation flag information is a first reference value. If the voice generation confidence information is less than or equal to the preset first threshold, the voice generation flag information is a third reference value. The first reference value, the second reference value and the preset first threshold can be set according to actual needs. For example, if the voice generation confidence information is greater than a preset 0.5, the voice generation flag information is 1. If the voice generation confidence information is less than or equal to 0.5, the voice generation flag information is 0.

[0104] Specifically, the object position information includes an x coordinate value and a y coordinate value of the object on the screen. In order to determine the position of the object on the screen, the screen is divided according to a preset size in this embodiment. Referring to FIG. 3, the screen is divided into 28*8 parts. The x coordinate value of the object close to the left side of the screen is 0, the x coordinate value of the object close to the right side of the screen is 1, the y coordinate value of the object close to the lower side of the screen is 0, and the y coordinate value of the object close to the upper side of the screen is 1.

[0105] In step S202, the number of objects in the current video frame is determined based on the object identification information.

[0106] The number of objects is the number of target objects in the current video frame. The target object is the object with the highest priority in the current video frame. For example, based on the object identification information, it is determined that the current video frame includes two persons, one machine and one animal, and the order of the objects from high to low priority is: person > machine > animal. It is determined that the target object is a person, and the number of objects in the current video frame is 2.

[0107] In step S203, if the number of objects is greater than or equal to a first number threshold and less than a second number threshold, the initial voice generation position information of the current video frame is determined based on the object position information.

[0108] Specifically, the first number threshold and the second number threshold are preset thresholds. The first number threshold and the second number threshold can be set according to actual needs. For example, the first number threshold is set to 1, and the second number threshold is set to 2. If the number of objects is greater than or equal to 1 and less than 2, that is, the number of objects is 1, the initial voice generation position information of the current video frame is determined based on the object position information.

[0109] The step of determining the initial sound emission position information of the current video frame based on the object position information specifically comprises: determining object position information of a target object from the object position information; and determining the initial sound emission position information of the current video frame based on the object position information of the target object. When determining the initial sound emission position information of the current video frame based on the object position information of the target object, the x-coordinate value in the object position information of the target object can be directly determined as the initial sound emission position information of the current video frame, or the initial sound emission position information of the current video frame can be determined based on the x-coordinate value in the object position information of the target object, which is not limited in the present application.

[0110] In step S204, the initial sound emission position information is adjusted based on the sound emission confidence information and the sound emission flag information to obtain target sound emission position information of the current video frame.

[0111] The target sound emission position information of the current video frame is a screen sound emission position determined based on the sound emission confidence information and the sound emission flag information. In the embodiment, the initial sound emission position information of the current video frame is determined based on the object position information, and then sound emission recognition and initial sound emission position information adjustment are performed based on the sound emission confidence information and the sound emission flag. Compared with the prior art of completely recognizing the sound emission of an object and then adjusting the sound emission position, the delay problem caused by the sound emission recognition time can be reduced, and the target sound emission position information obtained is more matched with the current video frame.

[0112] In a specific implementation, referring to FIG. 4, the step of obtaining the sound emission information of the current video frame in step S201 can include steps S301-S302, specifically as follows.

[0113] In step S301, an image recognition model is used to perform image recognition on the current video frame to obtain object information of an object contained in the current video frame.

[0114] The object information is information related to the object contained in the current video frame, for example, the object information includes a landmark point of an object, a landmark point of a human mouth, etc. The image recognition model is a neural network model pre-trained for recognizing the object information of the current video frame, and the image recognition model can be obtained by training a preset network model based on a sample video frame and real object information of the sample video frame.

[0115] The training process of the image recognition model comprises: inputting a sample video frame into a preset network model, outputting predicted object information of the sample video frame by the preset network model, determining a loss value based on the predicted object information, real object information and a loss function of the preset network model, training the preset network model based on a preset parameter learning rate if the loss value does not meet a preset condition, and continuing to perform the steps of inputting the sample video frame into the preset network model and outputting the predicted object information of the sample video frame by the preset network model until the loss value meets the preset condition. The trained preset network model is the image recognition model. The loss value meeting the preset condition can be that the loss value is less than a preset second threshold or a difference between loss values obtained at two adjacent times is less than a preset third threshold.

[0116] In step S302, the object information is converted to obtain sound emission information of the current video frame.

[0117] Specifically, the object information comprises object category, object position information in an image and object landmark point information. The object identification information can be determined by the object category, and the object position information in the screen can be determined based on the object position information in the image, thereby obtaining the object position information.

[0118] Further, the sound emission confidence information in the sound emission information can be obtained by the following steps: determining time sequence information of the object landmark points based on the object landmark point information; determining class variance data based on the time sequence information; and determining the sound emission confidence information based on the class variance data.

[0119] The sound emission flag information in the sound emission information can be obtained by the following steps: comparing the sound emission confidence information with a preset first threshold; if the sound emission confidence information is greater than the preset first threshold, determining the sound emission flag information as a first reference value; and if the sound emission confidence information is less than or equal to the preset first threshold, determining the sound emission flag information as a third reference value. The first reference value, the third reference value and the preset first threshold can be set according to actual needs. For example, the first reference value is set as 1, the third reference value is set as 0, and the preset first threshold is set as 0.5. If the sound emission confidence information is greater than 0.5, the sound emission flag information is determined as 1. If the sound emission confidence information is less than or equal to 0.5, the sound emission flag information is determined as 0.

[0120] In a specific implementation, with reference to FIG. 5, the step S202 of determining the number of objects in the current video frame based on the object identification information can comprise the following steps S401-S403, which are specifically as follows:

[0121] In step S401, a target object and an identification number of the target object are determined based on the object identification information and preset object priority information. The target object is an object with the highest priority among the objects included in the current video frame.

[0122] The object quantity of the current video frame mentioned in the foregoing step is the quantity of the object with the highest priority in the current video frame. To determine the object quantity of the current video frame, the embodiment pre-sets object priority information, for example, the object priority information is person > machine > musical instrument > animal. The target object is the object with the highest priority among the objects contained in the current video frame. The target object and the identification quantity of the target object can be determined based on the object identification information and the object priority information. For example, the object identification information is 00:1, 00:1, 00:2, and 00:4, and the priority of the person 00:1 is the highest in the object priority information. Therefore, based on the object identification information and the preset object priority information, it can be determined that the target object is a person, and the identification quantity of the target object is 2.

[0123] In step S402, the quantity of the target object is determined based on the identification quantity of the target object.

[0124] The quantity of the target object can be determined based on the identification quantity of the target object. For example, if the identification quantity of the target object is 2, the quantity of the target object is determined to be 2.

[0125] In step S403, the quantity of the target object is determined as the object quantity of the current video frame.

[0126] Specifically, after determining the quantity of the target object contained in the current video frame, the quantity of the target object can be determined as the object quantity of the current video frame. For example, if the quantity of the target object is determined to be 2, the object quantity of the current video frame is 2.

[0127] In a specific implementation, after determining the object quantity of the current video frame based on the object identification information in step S202, the method further includes: if the object quantity is less than a first quantity threshold or greater than or equal to a second quantity threshold, determining preset sound emission position information as initial sound emission position information of the current video frame; obtaining target sound emission position information of a historical video frame, and determining target sound emission position information of the current video frame based on the target sound emission position information of the historical video frame and the initial sound emission position information.

[0128] Specifically, the preset sound emission position information can be set according to actual needs. For example, referring to FIG. 3, the sound emission position is 13 / 14, and the sound effect is consistent, which is centered. Therefore, 14 is set as the preset sound emission position information. The historical video frame is z-1 video frames before the current video frame. z can be set according to actual needs. For example, z is 3. Of course, z can also be 1. When z is 1, the initial sound emission position information is directly determined as the target sound emission position information of the current video frame.

[0129] Further, when determining the target voice position information of the current video frame based on the target voice position information of the historical video frame and the initial voice position information, the average of the target voice position information of the historical video frame and the initial voice position information can be determined as the target voice position information of the current video frame, for example, the average of the initial voice position information and the target voice position information of two video frames before the current video frame can be determined as the target voice position information of the current video frame.

[0130] In a specific implementation, with reference to FIG. 6, the step S204 of adjusting the initial voice position information based on the voice confidence information and the voice flag information to obtain the target voice position information of the current video frame can include the following steps S501-S504, which are specifically as follows.

[0131] The step S501 includes performing voice recognition based on the voice confidence information and the voice flag information to obtain voice state information of a target object; the target object is an object with the highest priority among objects included in the current video frame.

[0132] Specifically, the voice state information of the target object indicates whether the target object is currently speaking, and the voice state information includes first state information and second state information, where the first state information indicates that the target object is currently speaking, and the second state information indicates that the target object is not currently speaking. In this embodiment, the voice recognition based on the voice confidence information and the voice flag information can improve the speed of voice recognition and reduce the delay problem caused by the voice recognition time.

[0133] In a specific embodiment, the step of performing voice recognition based on the voice confidence information and the voice flag information to obtain voice state information of a target object can specifically include the following steps: obtaining scene switching identification information, a preset confidence threshold, a first reference value, and a second reference value; if the voice confidence information is less than the preset confidence threshold, the voice flag information is not the first reference value, and the scene switching identification information is the second reference value, determining that the voice state information of the target object is the second voice state; if the voice confidence information is greater than or equal to the preset confidence threshold, and / or the voice flag information is the first reference value, and / or the scene switching identification information is not the second reference value, determining that the voice state information of the target object is the first voice state.

[0134] The scene switching identification information is used to represent whether the current video frame is in the process of scene switching. If the scene switching identification information is a second reference value, it represents that the current video frame is in the process of scene switching. If the scene switching identification information is not the second reference value, it represents that the current video frame is not in the process of scene switching. For example, the second reference value is 0. If the scene switching identification information is 0, it represents that the current video frame is in the process of scene switching. If the scene switching identification information is 1, it represents that the current video frame is not in the process of scene switching.

[0135] From the above process of voice recognition based on the voice confidence information and the voice flag information, it can be seen that the following three conditions need to be met simultaneously to determine that the voice state information of the target object is the second voice state: ① the voice confidence information is less than the pre-set confidence threshold; ② the voice flag information is not the first reference value; and ③ the scene switching identification information is the second reference value. If any one of the above three conditions is not met, it is determined that the voice state information of the target object is the first voice state. For example, ① the voice confidence information is less than 0.05; ② the voice flag is not 1; and ③ the scene switching flag information is 0. These three conditions are met simultaneously, and it is determined that the voice state information of the target object is the second voice state.

[0136] In step S502, if the voice state information is the first voice state and the target object is a person, the object position information of the target object is compared with the first position threshold and the second position threshold. The first voice state represents that the target object is speaking.

[0137] The first position threshold and the second position threshold are pre-set position thresholds for measuring whether the target object is in the middle of the screen. The first position threshold and the second position threshold can be set according to actual needs. For example, the first position threshold is set to 10 and the second position threshold is set to 17. If the voice state information is the first voice state and the target object is a person, the object position information of the target object is compared with the first position threshold and the second position threshold. Otherwise, if the voice state information is the first voice state but the target object is not a person, the target voice position information of the current video frame is determined based on the initial voice position information.

[0138] It should be noted that when determining the target voice position information of the current video frame based on the initial voice position information, the initial voice position information can be directly determined as the target voice position information of the current video frame, or the initial voice position information and the target voice position information of the historical video frame can be determined as the mean value of the target voice position information of the current video frame. The present application does not limit this.

[0139] In step S503, if the object position information of the target object is greater than the first position threshold and less than the second position threshold, the initial sound emission position information is adjusted based on the face deflection angle of the target object to obtain candidate sound emission position information.

[0140] Specifically, if the current video frame includes a person, the sound emission information further includes a face deflection angle, which represents a horizontal rotation angle of the face towards the audience. If the face is directly facing the audience, the face deflection angle is 0. If the face is facing the left side of the audience, the face deflection angle is positive. If the face is facing the right side of the audience, the face deflection angle is negative.

[0141] If the object position information of the target object is greater than the first position threshold and less than the second position threshold, indicating that the target object is in the middle of the screen, the initial sound emission position information is adjusted based on the face deflection angle of the target object to obtain candidate sound emission position information. For example, if the x-coordinate value in the object position information of the target object satisfies 10 < x < 17, the initial sound emission position information is adjusted based on the face deflection angle of the target object to obtain candidate sound emission position information.

[0142] In any case, if the object position information of the target object is less than or equal to the first position threshold or greater than or equal to the second position threshold, indicating that the target object is not in the middle of the screen, the initial sound emission position information is not adjusted, and the initial sound emission position information is directly determined as the candidate sound emission position information. For example, if the x-coordinate value in the object position information of the target object satisfies x ≥ 17 or x ≤ 10, the initial sound emission position information is not adjusted, and the initial sound emission position information is directly determined as the candidate sound emission position information.

[0143] In an embodiment, the step of adjusting the initial sound emission position information based on the face deflection angle of the target object to obtain candidate sound emission position information specifically includes: comparing the face deflection angle of the target object with a first angle threshold, a second angle threshold, a third angle threshold, and a fourth angle threshold; if the face deflection angle is greater than or equal to the first angle threshold and less than the second angle threshold, determining a first position information based on the face deflection angle and the first angle threshold; adding the first position information and a preset position information to obtain a second position information; and subtracting the initial sound emission position information and the second position information to obtain the candidate sound emission position information. In this case, the calculation process of the candidate sound emission position information can be represented as: wherein pos2 represents the candidate sound emission position information, pos1 represents the initial sound emission position information, A represents the first angle threshold, B represents the preset position information, and CEILING() represents the smallest integer greater than or equal to the given numerical expression.

[0144] The first angle threshold, the second angle threshold, the third angle threshold, the fourth angle threshold, and the preset position information can be set according to actual needs. For example, the first angle threshold is set to 15°, the second angle threshold is set to 60°, the third angle threshold is set to -60°, the fourth angle threshold is set to -15°, and the preset position information is set to 1. The calculation process of the candidate sound emission position information can be represented as:

[0145] In an embodiment, after comparing the face deflection angle of the target object with the first angle threshold, the second angle threshold, the third angle threshold, and the fourth angle threshold, the method further includes: if the face deflection angle is greater than the third angle threshold and less than or equal to the fourth angle threshold, determining the first position information based on the face deflection angle and the first angle threshold; subtracting the first position information from the preset position information to obtain the third position information; and subtracting the third position information from the initial sound emission position information to obtain the candidate sound emission position information. In this case, the calculation process of the candidate sound emission position information can be represented as: For example, the first angle threshold is 15°, the preset position information is 1, and if the face deflection angle angle of the target object satisfies -60° < angle < -15°, the calculation process of the candidate sound emission position information can be represented as:

[0146] In an embodiment, if the face deflection angle is greater than the fourth angle threshold and less than the first angle threshold, the initial sound emission position information is determined as the candidate sound emission position information. For example, if the face deflection angle angle satisfies -15° < angle < 15°, the determination process of the candidate sound emission position information can be represented as: pos2 = pos1, where pos2 represents the candidate sound emission position information, and pos1 represents the initial sound emission position information.

[0147] In step S504, the target sound emission position information of the current video frame is determined based on the candidate sound emission position information.

[0148] Specifically, when the target sound emission position information of the current video frame is determined based on the candidate sound emission position information, the candidate sound emission position information can be directly determined as the target sound emission position information of the current video frame, or the candidate sound emission position information and the average of the target sound emission position information of the historical video frames can be determined as the target sound emission position information of the current video frame, which is not limited in the present application.

[0149] Wherein, the historical video frame is z-1 video frames before the current video frame, and z can be set according to actual requirements, for example, if z is 3, the candidate speech position information and the average of the target speech position information of the two video frames before the current video frame are determined as the target speech position information of the current video frame. Of course, z can also be 1, and when z is 1, the candidate speech position information is directly determined as the target speech position information of the current video frame.

[0150] In a specific embodiment, after the speech recognition based on the speech confidence information and the speech flag information in the step S501 to obtain the speech state information of the target object, the method further includes: if the speech state information is the second speech state, updating the post-speech judgment queue value; judging whether the updated post-speech judgment queue value is greater than the queue upper limit value; if the updated post-speech judgment queue value is less than or equal to the queue upper limit value, obtaining the target speech position information of the historical video frame, and determining the target speech position information of the current video frame based on the target speech position information of the historical video frame and the initial speech position information; if the updated post-speech judgment queue value is greater than the queue upper limit value, determining the preset speech position information as the initial speech position information, and determining the target speech position information of the current video frame based on the target speech position information of the historical video frame and the initial speech position information. Wherein, the preset speech position information can be set according to actual requirements, for example, referring to FIG. 3, the sound effect is consistent when the speech position is 13 / 14, and both are centered, and 14 is set as the preset speech position information.

[0151] Further, the queue upper limit value can be set according to actual requirements, for example, the queue upper limit value is set to 9, and the updating of the post-speech judgment queue value can be increasing the post-speech judgment queue value by 1, for example, if the updated post-speech judgment queue value is less than or equal to 9, the average of the target speech position information of the historical video frame and the initial speech position information is determined as the target speech position information of the current video frame; if the updated post-speech judgment queue value is greater than 9, the preset speech position information is determined as the initial speech position information, and the average of the target speech position information of the historical video frame and the initial speech position information is determined as the target speech position information of the current video frame.

[0152] In a specific embodiment, after the updated post-speech judgment queue value is greater than the queue upper limit value, the method further includes: setting the post-speech judgment queue value to the sum of the queue upper limit and a preset fourth threshold value, for example, the preset fourth threshold value is 1, and the queue upper limit is 9, if the updated post-speech judgment queue value is greater than 9, the post-speech judgment queue value is set to 10, and such setting can prevent queue overflow caused by long time without appearing the person speaking.

[0153] In a specific embodiment, after the voice state information is the first voice state in the step S502, the method further includes: setting a post-voice judgment queue value, and setting the scene switching identification information as a fourth reference value. In this way, the scene switching identification information is set as the second reference value only in the scene switching process.

[0154] In a specific embodiment, after the object quantity is greater than or equal to the first quantity threshold and less than the second quantity threshold in the step S202, the method further includes: obtaining a difference value between the initial voice position information and the target voice position information of the previous video frame; and setting the scene switching identification information as the second reference value if the difference value is greater than a preset fifth threshold. In this way, whether the object included in the current video frame has abnormal jumping can be identified, and whether the scene switching occurs in the current video frame can be determined.

[0155] To better implement the voice position determination method in the embodiments of the present application, on the basis of the voice position determination method, a voice position determination device is further provided in the embodiments of the present application, as shown in FIG. 7, the voice position determination device 600 includes:

[0156] The information acquisition module 610 is configured to acquire voice information of a current video frame. The voice information includes object identification information, voice confidence information, voice flag information and object position information of an object included in the current video frame.

[0157] The quantity determination module 620 is configured to determine an object quantity of the current video frame based on the object identification information.

[0158] The position determination module 630 is configured to determine initial voice position information of the current video frame based on the object position information if the object quantity is greater than or equal to a first quantity threshold and less than a second quantity threshold.

[0159] The position adjustment module 640 is configured to adjust the initial voice position information based on the voice confidence information and the voice flag information to obtain target voice position information of the current video frame.

[0160] In the embodiments of the present application, when whether the object in the picture has voice is not identified, the voice position is determined based on the object position information first, and then the voice recognition and the initial voice position adjustment are performed based on the voice confidence information and the voice flag information. In this way, the voice position information that is more matched with the picture can be obtained, and the effect delay caused by the voice recognition time can be relieved.

[0161] In some embodiments of the present application, the information acquisition module 610 is specifically configured to:

[0162] perform image recognition on the current video frame by using an image recognition model to obtain object information of the object included in the current video frame.

[0163] The object information is converted to obtain sound information of the current video frame.

[0164] In some embodiments of the present application, the quantity determination module 620 is specifically configured to:

[0165] determine a target object and an identification quantity of the target object based on the object identification information and preset object priority information; the target object is an object with the highest priority among the objects included in the current video frame;

[0166] determine the quantity of the target object based on the identification quantity of the target object;

[0167] determine the quantity of the target object as the object quantity of the current video frame.

[0168] In some embodiments of the present application, after the quantity determination module 620 determines the object quantity of the current video frame based on the object identification information, the quantity determination module 620 is specifically further configured to:

[0169] if the object quantity is less than a first quantity threshold or greater than or equal to a second quantity threshold, determine the preset sound position information as initial sound position information of the current video frame;

[0170] obtain target sound position information of a historical video frame, and determine the target sound position information of the current video frame based on the target sound position information of the historical video frame and the initial sound position information.

[0171] In some embodiments of the present application, the position adjustment module 640 is specifically configured to:

[0172] perform sound recognition based on the sound confidence information and the sound flag information to obtain sound state information of the target object; the target object is an object with the highest priority among the objects included in the current video frame;

[0173] if the sound state information is a first sound state and the target object is a person, compare object position information of the target object with a first position threshold and a second position threshold; the first sound state indicates that the target object is currently making a sound;

[0174] if the object position information of the target object is greater than the first position threshold and less than the second position threshold, adjust the initial sound position information based on a face deflection angle of the target object to obtain candidate sound position information;

[0175] determine the target sound position information of the current video frame based on the candidate sound position information.

[0176] In some embodiments of the present application, the position adjustment module 640 is specifically further configured to:

[0177] compare the face deflection angle of the target object with the first angle threshold, the second angle threshold, the third angle threshold and the fourth angle threshold;

[0178] if the face deflection angle is greater than or equal to the first angle threshold and less than the second angle threshold, determining the first position information based on the face deflection angle and the first angle threshold;

[0179] adding the first position information and the preset position information to obtain the second position information;

[0180] subtracting the initial sound production position information and the second position information to obtain the candidate sound production position information.

[0181] In some embodiments of the present application, the position adjustment module 640 is specifically further used for:

[0182] if the face deflection angle is greater than the third angle threshold and less than or equal to the fourth angle threshold, determining the first position information based on the face deflection angle and the first angle threshold;

[0183] subtracting the first position information and the preset position information to obtain the third position information;

[0184] subtracting the initial sound production position information and the third position information to obtain the candidate sound production position information.

[0185] In some embodiments of the present application, the position adjustment module 640 is specifically further used for:

[0186] obtaining scene switching identification information, a preset confidence threshold, a first reference value and a second reference value;

[0187] if the sound production confidence information is less than the preset confidence threshold, the sound production flag bit information is not the first reference value and the scene switching identification information is the second reference value, determining that the sound production state information of the target object is the second sound production state; the second sound production state represents that the target object does not produce sound;

[0188] if the sound production confidence information is greater than or equal to the preset confidence threshold, and / or the sound production flag bit information is the first reference value, and / or the scene switching identification information is not the second reference value, determining that the sound production state information of the target object is the first sound production state.

[0189] In some embodiments of the present application, after the position adjustment module 640 obtains the sound production state information of the target object based on the sound production confidence information and the sound production flag bit information, the position adjustment module 640 is further used for:

[0190] if the sound production state information is the second sound production state, updating the post-sound judgment queue value;

[0191] determine whether the updated post-sound judgment queue value is greater than the queue upper limit value;

[0192] If the updated post-sound judgment queue value is less than or equal to the queue upper limit value, obtaining target sound position information of the historical video frame, and determining target sound position information of the current video frame based on the target sound position information of the historical video frame and the initial sound position information.

[0193] If the updated post-sound judgment queue value is greater than the queue upper limit value, determining preset sound position information as the initial sound position information, and determining target sound position information of the current video frame based on the target sound position information of the historical video frame and the initial sound position information.

[0194] The embodiment of the present application further provides a computer device which integrates any one of the sound position determination devices provided by the embodiment of the present application. The computer device comprises:

[0195] one or more processors;

[0196] a memory; and

[0197] one or more application programs, wherein the one or more application programs are stored in the memory and are configured to execute the steps in the sound position determination method in any one of the sound position determination method embodiments by the processor.

[0198] The embodiment of the present application further provides a computer device which integrates any one of the sound position determination devices provided by the embodiment of the present application. As shown in FIG. 8, it shows the structure schematic diagram of the computer device related to the embodiment of the present application, specifically:

[0199] The computer device can include a processor 801 with one or more processing cores, a memory 802 with one or more computer readable storage media, a power supply 803, and an input unit 804, and the like. Those skilled in the art can understand that the computer device structure shown in FIG. 8 does not constitute a limitation on the computer device, and can include more or fewer components than the illustration, or combine certain components, or different component arrangements. Among them:

[0200] The processor 801 is the control center of the computer device, connects the various parts of the computer device through various interfaces and lines, and performs various functions and processes data of the computer device by running or executing software programs and / or modules stored in the memory 802 and calling data stored in the memory 802, thereby overall monitoring the computer device. Optionally, the processor 801 can include one or more processing cores; preferably, the processor 801 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface and application program, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 801.

[0201] The memory 802 can be used to store software programs and modules, and the processor 801 executes various functions and data processing by running the software programs and modules stored in the memory 802. The memory 802 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc.; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 802 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. Accordingly, the memory 802 can also include a memory controller to provide access for the processor 801 to the memory 802.

[0202] The computer device further includes a power supply 803 for powering various components, and preferably the power supply 803 can be logically connected to the processor 801 through a power management system, so as to realize functions such as management of charging, discharging and power consumption management through the power management system. The power supply 803 can also include one or more than one direct current or alternating current power supply, a recharging system, a power failure detection circuit, a power converter or inverter, a power state indicator, etc. Any component.

[0203] The computer device can also include an input unit 804, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0204] Although not shown, the computer device can also include a display unit, etc., which will not be described here. Specifically, in the present embodiment, the processor 801 in the computer device will load the executable file corresponding to the process of one or more than one application program into the memory 802 according to the following instructions, and run the application program stored in the memory 802 by the processor 801, thereby realizing various functions, as follows:

[0205] acquire sound emission information of the current video frame; the sound emission information comprises object identification information, sound emission confidence information, sound emission flag information and object position information of an object included in the current video frame;

[0206] determine the number of objects in the current video frame based on the object identification information;

[0207] if the number of objects is greater than or equal to a first number threshold and less than a second number threshold, determine initial sound emission position information of the current video frame based on the object position information;

[0208] adjust the initial sound emission position information based on the sound emission confidence information and the sound emission flag information to obtain target sound emission position information of the current video frame.

[0209] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by related hardware controlled by the instructions, which can be stored in a computer readable storage medium and loaded and executed by a processor.

[0210] To this end, the embodiments of the present application provide a computer readable storage medium, which can include a read only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc. A computer program is stored on the storage medium and loaded by a processor to execute the steps in any sound emission position determination method provided by the embodiments of the present application. For example, the computer program loaded by the processor can execute the following steps:

[0211] acquire sound emission information of the current video frame; the sound emission information comprises object identification information, sound emission confidence information, sound emission flag information and object position information of an object included in the current video frame;

[0212] determine the number of objects in the current video frame based on the object identification information;

[0213] if the number of objects is greater than or equal to a first number threshold and less than a second number threshold, determine initial sound emission position information of the current video frame based on the object position information;

[0214] adjust the initial sound emission position information based on the sound emission confidence information and the sound emission flag information to obtain target sound emission position information of the current video frame.

[0215] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the detailed description of other embodiments above, which will not be repeated here.

[0216] In specific implementation, the above various units or structures can be implemented as independent entities, or can be combined as the same or several entities. The specific implementation of the above various units or structures can be referred to the method embodiments above, and will not be described here.

[0217] The specific implementation of the above various operations can be referred to the embodiments above, and will not be described here.

[0218] The method, device and computer equipment for determining a sound production position are described in detail above. The principle and implementation mode of the present application are described by using specific examples. The above embodiment is only used to help understand the method and core idea of the present application. Meanwhile, for those skilled in the art, the specific implementation mode and application range can be changed according to the idea of the present application. In summary, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A sound emitting position determination method, wherein, The method comprises: acquiring voice information of a current video frame; the voice information comprises object identification information, voice confidence information, voice flag information and object position information of an object included in the current video frame; determining the number of objects in the current video frame based on the object identification information; if the number of objects is greater than or equal to a first number threshold and less than a second number threshold, determining initial voice position information of the current video frame based on the object position information; adjusting the initial voice position information based on the voice confidence information and the voice flag information to obtain target voice position information of the current video frame.

2. The method of claim 1, wherein, The adjusting the initial voice position information based on the voice confidence information and the voice flag information to obtain target voice position information of the current video frame comprises: performing voice recognition based on the voice confidence information and the voice flag information to obtain voice state information of a target object; the target object is an object with the highest priority among the objects included in the current video frame; if the voice state information is a first voice state and the target object is a person, comparing object position information of the target object with a first position threshold and a second position threshold; the first voice state indicates that the target object is speaking; if the object position information of the target object is greater than the first position threshold and less than the second position threshold, adjusting the initial voice position information based on a face deflection angle of the target object to obtain candidate voice position information; determining target voice position information of the current video frame based on the candidate voice position information.

3. The method of claim 2, wherein, The adjusting the initial voice position information based on the face deflection angle of the target object to obtain candidate voice position information comprises: comparing the face deflection angle of the target object with a first angle threshold, a second angle threshold, a third angle threshold and a fourth angle threshold; if the face deflection angle is greater than or equal to the first angle threshold and less than the second angle threshold, determining first position information based on the face deflection angle and the first angle threshold; performing addition processing on the first position information and preset position information to obtain second position information; performing subtraction processing on the initial voice position information and the second position information to obtain candidate voice position information.

4. The method of claim 3, wherein, After the comparing the face deflection angle of the target object with a first angle threshold, a second angle threshold, a third angle threshold and a fourth angle threshold, the method further comprises: if the face deflection angle is greater than the third angle threshold and less than or equal to the fourth angle threshold, determining the first position information based on the face deflection angle and the first angle threshold; performing subtraction processing on the first position information and the preset position information to obtain third position information; performing subtraction processing on the initial voice position information and the third position information to obtain candidate voice position information.

5. The method of claim 2, wherein, The performing voice recognition based on the voice confidence information and the voice flag information to obtain voice state information of a target object comprises: Acquire scene switching identification information, pre-set confidence threshold, first reference value and second reference value; If the sound production confidence information is less than the pre-set confidence threshold, the sound production flag bit information is not the first reference value, and the scene switching identification information is the second reference value, it is determined that the sound production state information of the target object is the second sound production state; the second sound production state represents that the target object does not produce sound; If the sound production confidence information is greater than or equal to the pre-set confidence threshold, and / or the sound production flag bit information is the first reference value, and / or the scene switching identification information is not the second reference value, it is determined that the sound production state information of the target object is the first sound production state.

6. The method of claim 5, wherein, After the sound production recognition based on the sound production confidence information and the sound production flag bit information to obtain the sound production state information of the target object, the method further includes: If the sound production state information is the second sound production state, updating a post-sound judgment queue value; Determining whether the updated post-sound judgment queue value is greater than a queue upper limit value; If the updated post-sound judgment queue value is less than or equal to the queue upper limit value, acquiring target sound production position information of a historical video frame, and determining the target sound production position information of the current video frame based on the target sound production position information of the historical video frame and the initial sound production position information; If the updated post-sound judgment queue value is greater than the queue upper limit value, determining pre-set sound production position information as the initial sound production position information, and determining the target sound production position information of the current video frame based on the target sound production position information of the historical video frame and the initial sound production position information.

7. The method of claim 1, wherein, The method further includes: Determining a target object and an identification number of the target object based on the object identification information and pre-set object priority information; the target object is an object with the highest priority among the objects included in the current video frame; Determining the number of the target object based on the identification number of the target object; Determining the number of the target object as the number of objects in the current video frame.

8. The method of claim 1, wherein, After the number of objects in the current video frame is determined based on the object identification information, the method further includes: If the number of objects is less than the first number threshold or greater than or equal to the second number threshold, determining pre-set sound production position information as the initial sound production position information of the current video frame; Acquiring target sound production position information of a historical video frame, and determining the target sound production position information of the current video frame based on the target sound production position information of the historical video frame and the initial sound production position information.

9. The method of claim 1, wherein, The method further includes: Performing image recognition on the current video frame by using an image recognition model to obtain object information of the objects included in the current video frame; Converting the object information to obtain the sound production information of the current video frame.

10. The method of claim 1, wherein, The method further includes: Determining object position information of a target object from the object position information; Determining the initial sound production position information of the current video frame based on the object position information of the target object. Determine initial sound position information of the current video frame based on the object position information of the target object.

11. The method of claim 1, wherein, The sound confidence information in the sound information is obtained by the following steps: Determine time sequence information of the landmark points of the object based on the landmark point information of the object; Determine class variance data based on the time sequence information; Determine sound confidence information based on the class variance data.

12. The method of claim 1, wherein, The sound flag information in the sound information is obtained by the following steps: Compare the sound confidence information with a preset first threshold value; If the sound confidence information is greater than the preset first threshold value, determine the sound flag information as a first reference value; If the sound confidence information is less than or equal to the preset first threshold value, determine the sound flag information as a third reference value.

13. The method of claim 3, wherein, After comparing the face deflection angle of the target object with the first angle threshold value, the second angle threshold value, the third angle threshold value and the fourth angle threshold value, the method further includes: If the face deflection angle is greater than the fourth angle threshold value and less than the first angle threshold value, determine the initial sound position information as candidate sound position information.

14. The method of claim 6, wherein, After the updated rear sound judgment queue value is greater than the queue upper limit value, the method further includes: Set the rear sound judgment queue value as the sum of the queue upper limit and a preset fourth threshold value.

15. The method of claim 6, wherein, After the sound state information is the first sound state, the method further includes: A rear sound judgment queue value, and set the scene switching identification information as a fourth reference value.

16. The method of claim 5, wherein, After the object quantity is greater than or equal to the first quantity threshold value and less than the second quantity threshold value, the method further includes: Obtain the difference between the initial sound position information and target sound position information of a previous video frame; If the difference is greater than a preset fifth threshold value, set the scene switching identification information as a second reference value.

17. The method of claim 3, wherein, If the face deflection angle is greater than or equal to the first angle threshold and less than the second angle threshold, the calculation process of the candidate sound emitting position information is: wherein pos2 represents the candidate sound emitting position information, pos1 represents the initial sound emitting position information, A represents the first angle threshold, B represents the preset position information, and angle represents the face deflection angle.

18. The method of claim 4, wherein, If the face deflection angle is greater than the third angle threshold and less than or equal to the fourth angle threshold, the calculation process of the candidate sound production position information is: wherein pos2 represents the candidate sound production position information, pos1 represents the initial sound production position information, A represents the first angle threshold, B represents the preset position information, and angle represents the face deflection angle.

19. A sound-emitting position determination apparatus, wherein, The method includes: An information acquisition module is configured to acquire sound information of a current video frame; The sound information includes object identification information, sound confidence information, sound flag information and object position information of an object included in the current video frame; A quantity determination module is configured to determine an object quantity of the current video frame based on the object identification information; A position determination module is configured to determine initial sound position information of the current video frame based on the object position information if the object quantity is greater than or equal to a first quantity threshold value and less than a second quantity threshold value; A position adjustment module is configured to adjust the initial sound position information based on the sound confidence information and the sound flag information to obtain target sound position information of the current video frame.

20. A computer device, wherein, The computer device includes: One or more processors; Memory; and One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the processor to implement the sound position determination method in any one of claims 1 to 17.

Citation Information

Patent Citations

  • Sound production position determination method and device, computer equipment and storage medium

    CN118691672A

  • Sound source determination method and system, electronic equipment and readable storage medium

    CN116504272A

  • Techniques for object acquisition and tracking

    US20170186291A1

  • Data processing method, electronic apparatus, and storage medium

    US20240177335A1

  • Audio processing method and apparatus, electronic device, storage medium, and program product

    WO2024027315A1