An audio-video control method with intelligent tracking function of wireless microphone
Through the audio and video control system combining audio and video information, the problem of personnel position determination and the accuracy of intelligent microphone camera functions in video conferences is solved, real-time sound and position analysis in different environments is realized, and recognition efficiency and position inclusiveness are improved.
Patent Information
- Application Number
- CN202211324931.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-27
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-10-27
AI Technical Summary
In video conferencing, due to the irregular location of personnel and the different environments, it is difficult for the prior art to quickly and efficiently determine the location of each person present, and the camera function of the smart microphone cannot accurately find the video image of the speaker.
The audio information and video information in the space where the wireless microphone is located is obtained through the audio and video control system, and the audio attributes and matching audio information are used to distinguish between the audio and video information and the image information of the person, locate the person's position and send the position data to realize the amplification of audio and video surveillance information.
It realizes real-time analysis of the relationship between the voices made by people and the position of people in different environments, improves the efficiency of system identification of characters, and increases the inclusiveness of position coordinates.
Smart Images

Figure CN115695708B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio - video control, and specifically to an audio - video control method with a wireless microphone intelligent tracking function. Background Art
[0002] Currently, audio or video conference host systems mainly use microphones and speakers as carriers for sound signal transmission. Among them, a video conference refers to a meeting where people at two or more locations communicate face - to - face through communication devices and networks. According to the number of participating locations, video conferences can be divided into point - to - point conferences and multi - point conferences. In daily life, individuals who have no requirements for the security of conversation content, the quality of the conference, and the scale of the conference can use video software for video chatting. However, for business video conferences in government agencies, enterprises, and institutions, conditions such as a stable and secure network, reliable conference quality, and a formal conference environment are required, and professional video conference equipment needs to be used to form a dedicated video conference system. Since such video conference systems all use a television for display, they are also known as television conferences and video conferences.
[0003] However, when most enterprises hold video conferences, multiple groups of departments often hold them simultaneously, and there are multiple meeting participants in each group. During the meeting, people need to discuss with each other, which makes it difficult for the microphone to match the recorded audio information with the actual person during the sound pickup process. In addition, the spatial environments of the meetings are different, and the fixed positions of the people are also different, resulting in the inability to quickly and efficiently determine the positions of each person present, and the camera function of the intelligent microphone cannot accurately find the video image of the person who is speaking. Summary of the Invention
[0004] The purpose of the present invention is to provide an audio - video control method with a wireless microphone intelligent tracking function to solve the problems raised in the above background art.
[0005] To solve the above - mentioned technical problems, the present invention provides the following technical solution: An audio - video control method with a wireless microphone intelligent tracking function, including the following specific processes:
[0006] Step S100: The audio - video control system acquires the audio information and video information of the space where the wireless microphone is located. The audio information includes first audio information and second audio information, and the video information includes global person image information and local person image information.
[0007] Step S200: Based on the audio information in Step S100, parse the first audio information to obtain the first audio attributes, and distinguish audios with different attributes; according to the distinguished audio attributes, perform combined analysis with the global person image information, and match the specific person information in the global person image information corresponding to the audios with different attributes in the audio information.
[0008] Step S300: After the matching is completed, locate the positions of different personnel according to the second audio information and the local person image information, and send the position data of each person to the audio-video control system.
[0009] Step S400: When the audio-video control system receives the position data of all people, monitor whether the second audio information of the people corresponding to all the position data is updated. When the audio-video control system monitors a person with updated second audio information, send the position data of the person and globally magnify the local person image information corresponding to the person to obtain the audio-video monitoring information of the person.
[0010] Furthermore, the audio-video control system includes a debugging mode and a conference mode; the debugging mode is used to collect the first audio information and the global person image information, and the conference mode is used to collect the second audio information and the local person image information.
[0011] The debugging mode is used to place a wireless microphone. The wireless microphone is connected to the power supply of the audio-video conference host. The wireless microphone is provided with a power button for the microphone and shakes randomly up, down, left, and right when powered on to obtain the global person image information and the local person image information; when the audio-video control system analyzes a specific person, the camera on the wireless microphone will aim at the participant and obtain the position address of the person and send the position address to the audio-video control system; repeat the positioning, and the wireless microphone records the address of each person into the audio-video control system to complete the preliminary positioning of the system.
[0012] The conference mode is used to turn the camera to the participant corresponding to the position information and magnify the second audio information and the local person image information of the participant when the participant is speaking according to the position information already confirmed in the audio-video system.
[0013] Furthermore, the division of the first audio information and the second audio information includes the following process:
[0014] Step S110: The audio-video control system acquires the audio information in the audio acquisition stage, converts the audio information into digital signals, obtains the total time interval t0 between adjacent digital signals and the total information length p0 of the digital signals, and calculates the overall vocal fluctuation frequency index w = p0 / t0 of the digital signals; the time interval reflects two situations that may be included in the sound acquisition stage. One is that all sounds are chaotic and irregular, and the other is that the sounds are regularly emitted after the planning is completed; distinguishing these two situations is to distinguish the sound change rules applicable in many sound monitoring scenarios; and the vocal fluctuation frequency index represents the average change in the entire sound acquisition stage.
[0015] Step S120: Based on the vocal fluctuation frequency index in step S110, the audio-video control system traverses from the first digital signal in the sound acquisition stage, obtains the vocal fluctuation frequencies of the first digital signal and its adjacent digital signals, and calculates the frequency fluctuation difference by subtracting the vocal fluctuation frequency of the first digital signal from the overall vocal fluctuation frequency index; the frequency fluctuation difference represents the deviation degree between the sound fluctuation frequency and the average fluctuation frequency index in the monitoring scenario as time progresses; and there is a critical point for dividing the change of the sound, and the magnitude relationship between the frequency fluctuation differences before and after the critical point is not defined here, which can accommodate the sound change rules in more scenarios. For example, sometimes the frequency fluctuation in the pre-sound acquisition stage is large, and sometimes the frequency fluctuation in the pre-sound acquisition stage is small.
[0016] Step S130: Sequentially obtain the vocal fluctuation frequencies of adjacent digital signals in the sound acquisition stage and mark the turning digital signals. The turning digital signal is the digital signal corresponding to the quotient of the frequency fluctuation difference of the previous adjacent digital signal and the frequency fluctuation difference of the subsequent adjacent digital signal being negative, and the positive and negative of the frequency fluctuation differences of all digital signals after the turning digital signal are the same as those of the frequency fluctuation difference corresponding to the turning digital signal.
[0017] Step S140: Based on the judgment rule in step S130, the audio-video control system divides the audio information before the turning digital signal into the first audio information and the audio information after the turning digital signal into the second audio information.
[0018] Further, step S200 includes the following processes:
[0019] Step S210: Distinguish the digital signals corresponding to the first audio attributes according to the frequency similarity, classify the audio attributes corresponding to the digital signals with a frequency similarity greater than 95% into one category, and denote them as , j = {1, 2,... k}, j represents the number of different types of first audio attributes, represents the jth type of first audio attribute; and record the decibel characteristics of each audio as , where s is any natural number not equal to 0, and s represents the number of times the j-th first audio attribute appears during the sound collection phase. represents the decibel feature of the s-th occurrence of the j-th first audio attribute.
[0020] Frequency reflects the characteristics of sound, and everyone's voice frequency is different. First, divide the number of monitored people by frequency, and then analyze the decibel characteristics of each person separately, because the size of the decibel is affected by the distance between the receiving end and the generating end.
[0021] Step S220: Denote different first audio attributes and their corresponding decibel characteristics as a set A, and calculate the average decibel difference ratio of the decibel characteristics corresponding to the j-th first audio attribute in set A changing with time. , where represents the difference between two adjacent decibel characteristics in the j-th first audio attribute, and n represents the number of decibel characteristic differences, with n being at least 1; calculate the overall deviation index of different first audio attributes in set A. ;
[0022] Step S230: Classify the global images of different people in the global person image information to obtain the j-th global person image. Denote the person ratios of different global person images as a set B, and calculate the average person image ratio difference corresponding to the j-th global person image in set B changing with time. ; where represents the difference in person image ratios between two adjacent images in the j-th global person image, and m represents the number of person image ratio differences, with m being at least 1; calculate the overall deviation index of different global person images in set B. ;
[0023] The overall deviation index reflects the span of decibel magnitudes and the span of distances reflected by images among all monitored people. If these two spans are basically the same, it can indicate a connection between the change in decibels and the distance.
[0024] Step S240: Based on the overall deviation index Q in step S220 and the overall deviation index Z in step S230, calculate the deviation index similarity. , if the deviation index similarity is greater than the similarity threshold, it indicates that the change in the decibels of the personnel is related to the movement of the positions of the personnel in the global image, and perform a one-to-one correspondence with a similarity greater than 99% between the average decibel difference ratio in set A and the average person image ratio difference in set B to obtain the audio attributes corresponding to different people in the global person image information. When explaining the relevance between the change in decibels and the distance, further analyze the decibel change law of each type of audio attribute and the change law of the global image, and after the correspondence, it is possible to match which person emits which frequency.
[0025] Further, step S300 includes the following process:
[0026] Based on the data after one-to-one correspondence in step S240, taking the person who first makes a sound in the second audio information as the starting person, obtaining the person ratios of all person images in the local person image information, sorting the person ratios from largest to smallest, and setting the position coordinates corresponding to the image with the smallest person ratio as the starting coordinates;
[0027] When any person makes a sound during monitoring, based on the size relationship between the person ratio of the person making a sound at this time in the local person image and the person ratio of the starting person, perform sector adaptation on the starting coordinates to obtain the position coordinates of the second person; Sector adaptation means taking the starting coordinates as the center of the sector, converting the relationship between the person ratio of the person making a sound and the person ratio of the starting person into mathematical data, and performing co-directional estimation with the data as the radius. Because during the monitoring process of the wireless microphone, there are various possible positions for the distribution of people in a stable state, such as matrix distribution, circular distribution, etc., and when obtaining person images, first obtain the person images of all people in one camera position, and the person ratio in the image of each person can be obtained. Because the positions of each person are different, the person image ratios will also be different when the image acquisition position remains unchanged. Using sector adaptation can increase the inclusiveness of the position coordinates.
[0028] Further, the audio-visual control method includes an audio-visual control system, and the audio-visual control system includes a space information acquisition module, a space information analysis module, a position data acquisition module, and a monitoring data amplification module;
[0029] The space information acquisition module is used to acquire the data information of the space where the wireless microphone is located and transmit the data information to the space information analysis module; the space information analysis module is used to analyze the data information from the space information acquisition module; the position data acquisition module is used to judge the position information of the people in the space according to the data information after the analysis is completed; the monitoring data amplification module is used to globally amplify the people in the new data information to obtain the audio-visual monitoring information of the corresponding people when new data information is added to the space information acquisition module.
[0030] Further, the space information acquisition module includes an audio information acquisition module and a video information acquisition module; the audio information acquisition module acquires audio information, and the audio information includes first audio information and second audio information; the video information acquisition module acquires video information, and the video information acquisition module includes a global person image information acquisition module and a local person image information acquisition module;
[0031] The audio information acquisition module includes a digital signal conversion module, a sound wave frequency index calculation module, a turning digital signal marking module, and an audio information division module;
[0032] The digital signal conversion module is used to convert audio information into digital signals. The vocal fluctuation frequency index calculation module obtains the total time interval t0 between adjacent digital signals and the total information length p0 of the digital signals, and calculates the overall vocal fluctuation frequency index w = p0 / t0 of the digital signals.
[0033] The turning digital signal marking module traverses the vocal fluctuation frequencies of the first digital signal and adjacent digital signals, and obtains the frequency fluctuation difference by taking the difference between the vocal fluctuation frequency of the first digital signal and the overall vocal fluctuation frequency index. Sequentially obtain the vocal fluctuation frequencies of adjacent digital signals in the sound acquisition stage and mark the turning digital signals. The turning digital signal is the digital signal corresponding to the quotient of the frequency fluctuation difference between the previous adjacent digital signal and the frequency fluctuation difference between the subsequent adjacent digital signal being negative, and the positive and negative of the frequency fluctuation differences of all digital signals after the turning digital signal are the same as the positive and negative of the frequency fluctuation difference corresponding to the turning digital signal.
[0034] The audio information division module is used to divide the audio information before the turning digital signal into first audio information, and the audio information after the turning digital signal into second audio information.
[0035] Furthermore, the spatial information analysis module includes an audio information analysis module, a video information analysis module, and a person matching module; the audio information analysis module includes an audio attribute classification module, an average decibel difference ratio calculation module, and an audio attribute deviation index calculation module; the video information analysis module includes a global person image classification module, a person image ratio difference calculation module, and a global person image deviation index calculation module; the person matching module includes a deviation index similarity calculation module and a person audio attribute correspondence module;
[0036] The audio attribute analysis module classifies different audio attributes. The average decibel difference ratio calculation module is used to record the decibel characteristics of each audio and record different first audio attributes and their corresponding decibel characteristics as a set A, and calculate the average decibel difference ratio of the decibel characteristics corresponding to the j-th first audio attribute in the set A changing with time; the audio attribute deviation index calculation module is used to calculate the overall deviation index of different first audio attributes in the set A.
[0037] The global person image classification module is used to classify the person images in the global person image information acquisition module, and record the person image ratios corresponding to different person images as a set B; the person image ratio difference calculation module is used to calculate the average person image ratio difference of different global person images in the set B changing with time, and the global person image deviation index calculation module is used to calculate the overall deviation index of different global person images.
[0038] The deviation index similarity calculation module is used to compare the numerical similarity in the global person image deviation index calculation module and the audio attribute deviation index calculation module. When it is greater than the similarity threshold, it indicates that the change in the person's decibel is related to the movement of the person's position in the global image; the person audio attribute corresponding module performs a one-to-one correspondence with a similarity greater than 99% on the average decibel difference ratio in set A and the average person image ratio difference in set B to obtain the audio attributes corresponding to different people in the global person image information.
[0039] Furthermore, the position data acquisition module includes a person image ratio sorting module, a starting coordinate setting module, and a sector adaptation module;
[0040] The person image ratio sorting module is used to take the person who first makes a sound in the second audio information as the starting person, and obtain the person ratios of all person images in the local person image information, and sort the person ratios from largest to smallest; the starting coordinate setting module sets the position coordinate corresponding to the person image with the smallest person ratio as the starting coordinate;
[0041] The sector adaptation module is used to, when any person in the monitoring makes a sound, based on the relationship between the person ratio of the person making the sound at this time in the local person image and the person ratio of the starting person, perform sector adaptation on the starting coordinate to obtain the position coordinate of the second person.
[0042] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: The present invention solves the problem that in the scenario applicable to the intelligent tracking of wireless microphones, due to the complexity of personnel, it is impossible to simply and efficiently distinguish the source of the sound from the specific correspondence of the person, so that when monitoring, the sound is captured but the position of the sound cannot be determined. Moreover, the present invention adapts to all environments of microphone monitoring for adjustable positioning, uses the method of combining audio information and video information for analysis to judge the relevance of the spatial sound position, and then obtains the correspondence between the sound and the person. The present invention enables the real-time relationship between the sound emitted by a person and the person's position to be distinguished in different environments, improving the efficiency of the system in identifying people; and the present invention increases the inclusiveness of the scene where the position coordinate is located by determining the coordinate through the image ratio. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0044] Figure 1 is the system structure diagram of an audio-video control method with the intelligent tracking function of a wireless microphone according to the present invention;
[0045] Figure 2 is the step diagram of an audio-video control method with the intelligent tracking function of a wireless microphone according to the present invention.
[0046] Figure 3 This is the block diagram of the audio - video control principle of an audio - video control method with a wireless microphone intelligent tracking function according to the present invention;
[0047] Figure 4 This is the block diagram of the conference microphone principle of an audio - video control method with a wireless microphone intelligent tracking function according to the present invention;
[0048] Figure 5 This is the block diagram of the host of the camera device of an audio - video control method with a wireless microphone intelligent tracking function according to the present invention. Specific embodiments
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] Please refer to Figures 1-5 , the present invention provides a technical solution: an audio - video control method with a wireless microphone intelligent tracking function, including the following specific processes:
[0051] Step S100: The audio - video control system acquires the audio information and video information in the space where the wireless microphone is located. The audio information includes the first audio information and the second audio information, and the video information includes the global human image information and the local human image information;
[0052] Step S200: Based on the audio information in step S100, the first audio information is analyzed to obtain the first audio attribute, and the audio with different attributes is distinguished; according to the distinguished audio attributes and the global human image information, a combined analysis is performed to match the specific human information in the global human image information corresponding to the audio with different attributes in the audio information;
[0053] Step S300: After the matching is completed, the positions of different people are located according to the second audio information and the local human image information, and the position data of each person is sent to the audio - video control system;
[0054] Step S400: When the audio - video control system receives the position data of all people, it monitors whether the second audio information of the people corresponding to all the position data is updated. When the audio - video control system monitors the person with updated second audio information, it sends the position data of the person and globally magnifies the local human image information corresponding to the person to obtain the audio - video monitoring information of the corresponding person.
[0055] As in the embodiment Figure 3 As shown, the audio - video control system includes a camera device host and multiple conference microphones. Each conference microphone is connected to the camera device host by wireless communication;
[0056] As Figure 4 As shown, the conference microphone includes a microphone control main chip, a microphone input circuit, a first 2.4G transceiver circuit, a key circuit, and a microphone power supply circuit. The microphone input circuit, the first 2.4G transceiver circuit, and the key circuit are respectively connected to the microphone control main chip, and the microphone power supply circuit supplies power to the microphone control main chip, the microphone input circuit, the first 2.4G transceiver circuit, and the key circuit;
[0057] As Figure 5 As shown, the camera device host has a camera and a speaker; the circuit structure of the camera device host includes a camera device host chip, an audio output circuit, a host key, a second 2.4G transceiver circuit, a speaker output circuit, a USB interface circuit, a camera, a track control circuit, and a power supply circuit; the camera device host chip is connected to the track control circuit, and the track control circuit is connected to the camera; the camera device host chip is electrically connected to the USB interface circuit at the same time, and the USB interface circuit is connected to the camera at the same time; the camera device host chip is electrically connected to the audio output circuit, and this audio output circuit is connected to an audio communication device; the camera device host chip is electrically connected to the speaker output circuit, and this speaker output circuit is used to connect to the speaker; the camera device host chip is also electrically connected to the second 2.4G transceiver circuit;
[0058] The microphone control main chip and the camera device host chip are Bluetooth 5.3 LE Audio chips;
[0059] The audio - video control system includes a debugging mode and a conference mode; the debugging mode is used to collect first audio information and global person image information, and the conference mode is used to collect second audio information and local person image information;
[0060] The debugging mode is used to place a wireless microphone. The wireless microphone turns on the power of the audio - video conference host. The wireless microphone is provided with a power button for the microphone and shakes randomly up, down, left, and right when the power is turned on to obtain global person image information and local person image information; when the audio - video control system analyzes a specific person, the camera on the wireless microphone will aim at the participant and obtain the location address of the person and send the location address to the audio - video control system; repeat the positioning, and the wireless microphone records the address of each person into the audio - video control system to complete the preliminary positioning of the system;
[0061] The conference mode is used to turn the camera to the participant corresponding to the location information and magnify the second audio information and local person image information of the participant according to the location information already confirmed in the audio - video system when the participant is speaking.
[0062] For example, in an embodiment: The main body of the imaging device has XYZ three axes, and drives the camera through the XYZ three axes;
[0063] After the main body of the imaging device starts to connect to the camera, the XYZ axis track performs a scan. When the camera faces the direction of the conference microphone, the conference microphone will send a command to the main body of the imaging device at this time, and the main body of the imaging device will record the position of the XYZ axis track corresponding to this conference microphone at this time;
[0064] The conference microphone will send a set of control codes to the main body of the imaging device using a wireless signal. After receiving the control codes, the main body of the imaging device will control the camera to turn to the direction of this conference microphone according to the position data of the XYZ axis track of this conference microphone;
[0065] The conference microphone encodes and decodes the received audio signal through the LE Audio LC3 and LC3+ codec technologies, and wirelessly transmits the formed voice packet to the main body of the imaging device using the TDMA multiple access technology.
[0066] The division of the first audio information and the second audio information includes the following process:
[0067] Step S110: The audio-video control system acquires the audio information in the audio acquisition stage, converts the audio information into a digital signal, obtains the total time interval t0 between adjacent digital signals and the total information length p0 of the digital signal, and calculates the overall sound fluctuation frequency index w = p0 / t0; The time interval reflects two situations that may be included in the sound acquisition stage. One is that all sounds are chaotic and irregular, and the other is that the sound is regular after the planning is completed; Distinguishing these two situations is to distinguish the sound change rules applicable in many sound monitoring scenarios; And the sound fluctuation frequency index represents the average change in the entire sound acquisition stage;
[0068] Step S120: Based on the sound fluctuation frequency index in step S110, the audio-video control system traverses from the first digital signal in the sound acquisition stage, obtains the sound fluctuation frequencies of the first digital signal and the adjacent digital signals, and calculates the difference between the sound fluctuation frequency of the first digital signal and the overall sound fluctuation frequency index to obtain the frequency fluctuation difference; The frequency fluctuation difference represents the deviation degree of the sound fluctuation frequency from the average fluctuation frequency index in the monitoring scenario as time progresses; And there will be a critical point for the change of the sound, and the size relationship between the frequency fluctuation differences before and after the critical point is not defined here, which can accommodate the sound change rules in more scenarios, such as: sometimes the frequency fluctuation is large in the pre-sound acquisition stage, and sometimes the frequency fluctuation is small in the pre-sound acquisition stage;
[0069] Step S130: Sequentially obtain the vocal fluctuation frequencies of adjacent digital signals in the sound acquisition stage and mark the turning digital signals. The turning digital signal is the digital signal corresponding to the quotient of the frequency fluctuation difference between the previous adjacent digital signal and the frequency fluctuation difference between the subsequent adjacent digital signal being negative, and the positive and negative of the frequency fluctuation differences of all digital signals after the turning digital signal are the same as those of the frequency fluctuation difference corresponding to the turning digital signal.
[0070] Step S140: Based on the judgment rule in Step S130, the audio - video control system divides the audio information before the turning digital signal into the first audio information and the audio information after the turning digital signal into the second audio information.
[0071] Step S200 includes the following processes:
[0072] Step S210: Distinguish the digital signals corresponding to the first audio attributes by frequency similarity. Group the digital signals with a frequency similarity greater than 95% into one category and denote it as , j = {1, 2,... k}, where j represents the number of different types of the first audio attributes, represents the j - th type of the first audio attribute; and record the decibel characteristics of each audio as , s is any non - zero natural number, s represents the number of times the j - th type of the first audio attribute appears in the sound acquisition stage, represents the decibel characteristic of the s - th appearance of the j - th type of the first audio attribute;
[0073] Frequency reflects the characteristics of sound, and everyone's voice frequency is different. First, divide the number of people monitored by frequency, and then analyze the decibel characteristics of each person separately, because the size of the decibel is affected by the distance between the receiving end and the generating end.
[0074] Step S220: Denote different types of the first audio attributes and their corresponding decibel characteristics as a set A, and calculate the average decibel difference ratio of the decibel characteristics corresponding to the j - th type of the first audio attribute in the set A changing with time , where represents the difference between two adjacent decibel characteristics in the j - th type of the first audio attribute, and n represents the number of decibel characteristic differences, n is at least 1; calculate the overall deviation index of different types of the first audio attributes in the set A ;
[0075] Step S230: Classify the global images of different people in the global person image information to obtain the j - th type of global person image , denote the person ratios of different types of global person images as a set B, and calculate the average person image ratio difference corresponding to the j - th type of global person image changing with time ; where denotes the difference in the proportion of the person images between two adjacent images in the j-th type of global person image, m represents the number of differences in the proportion of the person images, and m is at least 1; calculate the overall deviation index of different types of global person images in set B ;
[0076] The overall deviation index reflects the span of the decibel levels and the span of the image-reflected distances among all the monitored people. If these two spans are basically the same, it can indicate that there is a connection between the change in decibels and the distance;
[0077] Step S240: Based on the overall deviation index Q in step S220 and the overall deviation index Z in step S230, calculate the similarity of the deviation index , if the similarity of the deviation index is greater than the similarity threshold, it indicates that the change in the decibels of the person is related to the movement of the person's position in the global image, and a one-to-one correspondence with a similarity greater than 99% is performed between the average decibel difference ratio in set A and the average difference in the proportion of the person images in set B to obtain the audio attributes corresponding to different people in the global person image information. When explaining that the change in decibels is related to the distance, further analyze the change rules of the decibels of each type of audio attribute and the change rules of the global image, and after the correspondence, it is possible to match which person emits which frequency.
[0078] For example: There are 3 types of first audio attributes on-site , , , and each type of audio attribute contains three decibel characteristics, corresponding to , then set A is expressed as { }, , , ; then ;
[0079] There are 3 types of global person images, then set B is { , }, then %, 7%, 7%, ;
[0080] The process of one-to-one correspondence is to compare the similarities between , { }, { } in set A and { }, { }, { } in set B respectively and sequentially.
[0081] Step S300 includes the following process:
[0082] Based on the data after one-to-one correspondence in step S240, taking the person who makes the first sound in the second audio information as the starting person, obtaining the person ratios of all person images in the local person image information, sorting the person ratios from largest to smallest, and setting the position coordinates corresponding to the image with the smallest person ratio as the starting coordinates;
[0083] When any person makes a sound during monitoring, based on the size relationship between the person ratio of the person making the sound at this time in the local person image and the person ratio of the starting person, perform sector adaptation on the starting coordinates to obtain the position coordinates of the second person; Sector adaptation means taking the starting coordinates as the center of the sector, converting the relationship between the person ratio of the person making the sound and the person ratio of the starting person into mathematical data, and performing co-directional estimation with the data as the radius. Because during the monitoring of wireless microphones, there are various possible positions for the distribution of people in a stable state, such as matrix distribution, circular distribution, etc., and when obtaining person images, first obtain the person images of all people in one camera position, and the person ratio relationship of each person in the image can be obtained. Because the positions of each person are different, the person image ratios will also be different when the image acquisition position remains unchanged. Using sector adaptation can increase the inclusiveness of the position coordinates.
[0084] The audio-video control method includes an audio-video control system, and the audio-video control system includes a space information acquisition module, a space information analysis module, a position data acquisition module, and a monitoring data amplification module;
[0085] The space information acquisition module is used to acquire the data information of the space where the wireless microphone is located and transmit the data information to the space information analysis module; the space information analysis module is used to analyze the data information from the space information acquisition module; the position data acquisition module is used to judge the position information of the people in the space according to the data information after analysis; the monitoring data amplification module is used to globally amplify the people in the new data information to obtain the audio-video monitoring information of the corresponding people when new data information is added to the space information acquisition module.
[0086] The space information acquisition module includes an audio information acquisition module and a video information acquisition module; the audio information acquisition module acquires audio information, and the audio information includes first audio information and second audio information; the video information acquisition module acquires video information, and the video information acquisition module includes a global person image information acquisition module and a local person image information acquisition module;
[0087] The audio information acquisition module includes a digital signal conversion module, a sound wave frequency index calculation module, a turning digital signal marking module, and an audio information division module;
[0088] The digital signal conversion module is used to convert audio information into digital signals. The vocal fluctuation frequency index calculation module obtains the total time interval t0 between adjacent digital signals and the total information length p0 of the digital signals, and calculates the overall vocal fluctuation frequency index w = p0 / t0 of the digital signals;
[0089] The turning digital signal marking module traverses the vocal fluctuation frequencies of the first digital signal and adjacent digital signals, and obtains the frequency fluctuation difference by taking the difference between the vocal fluctuation frequency of the first digital signal and the overall vocal fluctuation frequency index. Sequentially obtain the vocal fluctuation frequencies of adjacent digital signals in the sound acquisition stage and mark the turning digital signals. The turning digital signal is the digital signal corresponding to the quotient of the frequency fluctuation difference between the previous adjacent digital signal and the frequency fluctuation difference between the subsequent adjacent digital signal being negative, and the positive and negative of the frequency fluctuation differences of all digital signals after the turning digital signal are the same as the positive and negative of the frequency fluctuation difference corresponding to the turning digital signal;
[0090] The audio information division module is used to divide the audio information before the turning digital signal into first audio information, and the audio information after the turning digital signal into second audio information.
[0091] The spatial information analysis module includes an audio information analysis module, a video information analysis module, and a person matching module; the audio information analysis module includes an audio attribute classification module, an average decibel difference ratio calculation module, and an audio attribute deviation index calculation module; the video information analysis module includes a global person image classification module, a person image ratio difference calculation module, and a global person image deviation index calculation module; the person matching module includes a deviation index similarity calculation module and a person audio attribute correspondence module;
[0092] The audio attribute analysis module classifies different audio attributes. The average decibel difference ratio calculation module is used to record the decibel characteristics of each audio and record different first audio attributes and their corresponding decibel characteristics as a set A, and calculate the average decibel difference ratio of the decibel characteristics corresponding to the j-th first audio attribute in the set A changing with time; the audio attribute deviation index calculation module is used to calculate the overall deviation index of different first audio attributes in the set A;
[0093] The global person image classification module is used to classify the person images in the global person image information acquisition module, and record the person image ratios corresponding to different person images as a set B; the person image ratio difference calculation module is used to calculate the average person image ratio difference of different global person images in the set B changing with time, and the global person image deviation index calculation module is used to calculate the overall deviation index of different global person images;
[0094] The deviation index similarity calculation module is used to compare the numerical similarity in the global person image deviation index calculation module and the audio attribute deviation index calculation module. When it is greater than the similarity threshold, it indicates that the change in the person's decibel is related to the movement of the person's position in the global image; the person audio attribute corresponding module performs a one-to-one correspondence with a similarity greater than 99% between the average decibel difference ratio in set A and the average person image ratio difference in set B to obtain the audio attributes corresponding to different people in the global person image information.
[0095] The position data acquisition module includes a person image ratio sorting module, a starting coordinate setting module, and a sector adaptation module;
[0096] The person image ratio sorting module is used to take the person who first makes a sound in the second audio information as the starting person, and obtain the person ratios of all person images in the local person image information, and sort the person ratios from largest to smallest; the starting coordinate setting module sets the position coordinate corresponding to the person image with the smallest person ratio as the starting coordinate;
[0097] The sector adaptation module is used to, when any person makes a sound during monitoring, based on the relationship between the person ratio of the person making the sound at this time in the local person image and the person ratio of the starting person, perform sector adaptation on the starting coordinate to obtain the position coordinate of the second person.
[0098] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.
[0099] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An audio - video control method with intelligent tracking function of wireless microphone, characterized in that, it includes the following specific processes: Step S100: The audio - video control system acquires the audio information and video information in the space where the wireless microphone is located. The audio information includes the first audio information and the second audio information, and the video information includes the global person image information and the local person image information; Step S200: Based on the audio information in Step S100, parse the first audio information to obtain the first audio attribute, and distinguish the audio with different attributes; According to the combined analysis of the distinguished audio attributes and the global person image information, match the specific person information in the global person image information corresponding to the audio with different attributes in the audio information; Step S300: After the matching is completed, locate the positions of different people according to the second audio information and the local person image information, and send the position data of each person to the audio - video control system; Step S400: When the audio - video control system receives the position data of all people, monitor whether the second audio information of the people corresponding to all position data is updated. When the audio - video control system monitors the person with updated second audio information, send the position data of the person and globally magnify the local person image information corresponding to the person to obtain the audio - video monitoring information of the person.
2. The audio - video control method with intelligent tracking function of wireless microphone according to claim 1, characterized in that: The audio - video control system includes a debugging mode and a meeting mode; the debugging mode is used to collect the first audio information and the global person image information, and the meeting mode is used to collect the second audio information and the local person image information; The debugging mode is used to place the wireless microphone. The wireless microphone is powered on to the audio - video conference host. The wireless microphone is provided with a power button of the microphone and shakes randomly up, down, left, and right without a target to obtain the global person image information and the local person image information when powered on; when the audio - video control system analyzes a specific person, the camera on the wireless microphone will aim at the conference participants, obtain the position address of the person, and send the position address to the audio - video control system; repeat the positioning, and the wireless microphone records the address of each person into the audio - video control system to complete the preliminary positioning of the system; The meeting mode is used to turn the camera to the conference participant corresponding to the position information and magnify the second audio information and the local person image information of the conference participant according to the position information already confirmed in the audio - video system when the conference participant speaks.
3. The audio - video control method with intelligent tracking function of wireless microphone according to claim 1, characterized in that: The division of the first audio information and the second audio information includes the following process: Step S110: The audio - video control system acquires the audio information in the audio acquisition stage, converts the audio information into a digital signal, obtains the total time interval t0 between adjacent digital signals and the total information length p0 of the digital signal, and calculates the overall sound - generating fluctuation frequency index w = p0 / t0; Step S120: Based on the voice fluctuation frequency index in Step S110, the audio-video control system traverses from the first digital signal in the sound acquisition stage, obtains the voice fluctuation frequencies of the first digital signal and adjacent digital signals, and calculates the frequency fluctuation difference by subtracting the voice fluctuation frequency of the first digital signal from the overall voice fluctuation frequency index; Step S130: Sequentially obtain the voice fluctuation frequencies of adjacent digital signals in the sound acquisition stage and mark the turning digital signals. The turning digital signal is the digital signal corresponding to the quotient of the frequency fluctuation difference between the previous adjacent digital signal and the next adjacent digital signal being negative, and the positive and negative of the frequency fluctuation differences of all digital signals after the turning digital signal are the same as those of the frequency fluctuation difference corresponding to the turning digital signal; Step S140: Based on the judgment rule in Step S130, the audio-video control system divides the audio information before the turning digital signal into first audio information and the audio information after the turning digital signal into second audio information.
4. An audio-video control method with a wireless microphone intelligent tracking function according to claim 2, wherein: The process of Step S200 includes the following: Step S210: Distinguish the digital signals corresponding to the first audio attributes according to frequency similarity. Group the audio attributes corresponding to the digital signals with a frequency similarity greater than 95% into one category and denote it as , where j = {1, 2,... k}, and j represents the number of different types of first audio attributes, represents the j-th type of first audio attribute; and record the decibel feature of each audio as , where s is any natural number not equal to 0, and s represents the number of times the j-th type of first audio attribute appears in the sound acquisition stage, represents the decibel feature of the s-th appearance of the j-th type of first audio attribute; Step S220: Denote different types of first audio attributes and their corresponding decibel characteristics as a set A, and calculate the average decibel difference ratio of the decibel characteristics corresponding to the j-th type of first audio attribute in the set A changing with time , where represents the difference between two adjacent decibel characteristics in the j-th type of first audio attribute, n represents the number of decibel characteristic differences, and n is at least 1; calculate the overall deviation index of different types of first audio attributes in the set A ; Step S230: Classify the global images of different people in the global person image information to obtain the j-th type of global person image , denote the person ratios of different types of global person images as a set B, and calculate the average person image ratio difference corresponding to the j-th type of global person image in set B over time ; where represents the person image ratio difference between two adjacent images in the j-th type of global person image, m represents the number of person image ratio differences, and m is at least 1; calculate the overall deviation index of different types of global person images in set B ; Step S240: Calculate the deviation index similarity based on the overall deviation index Q in Step S220 and the overall deviation index Z in Step S230 , if the deviation index similarity is greater than the similarity threshold, it indicates that the change in the personnel decibel is related to the movement of the personnel position in the global image, and a one-to-one correspondence with a similarity greater than 99% is performed on the average decibel difference ratio in set A and the average human image ratio difference in set B to obtain the audio attributes corresponding to different people in the global human image information.
5. An audio-video control method with a wireless microphone intelligent tracking function according to claim 3, wherein: The process of Step S300 includes the following: Based on the data corresponding one by one in Step S240, taking the person who first makes a sound in the second audio information as the starting person, obtaining the person ratios of all person images in the local person image information, sorting the person ratios from largest to smallest, and setting the position coordinates corresponding to the image with the smallest person ratio as the starting coordinates; When any person makes a sound during monitoring, based on the size relationship between the person ratio of the person making a sound at this time in the local person image and the person ratio of the starting person, perform sector adaptation on the starting coordinates to obtain the position coordinates of the second person; the sector adaptation means taking the starting coordinates as the center of the sector, converting the relationship between the person ratio of the person making a sound and the person ratio of the starting person into mathematical data, and performing co-directional estimation with the data as the radius.
6. An audio-video control method with a wireless microphone intelligent tracking function according to claim 1, wherein: The audio-video control method includes an audio-video control system, and the audio-video control system includes a space information acquisition module, a space information analysis module, a position data acquisition module, and a monitoring data amplification module; The space information acquisition module is used to acquire the data information of the space where the wireless microphone is located and transmit the data information to the space information analysis module; the space information analysis module is used to analyze the data information from the space information acquisition module; the position data acquisition module is used to judge the position information of the personnel in the space according to the data information after the analysis is completed; the monitoring data amplification module is used to globally amplify the personnel in the new data information to obtain the audio-video monitoring information of the corresponding personnel when new data information is added to the space information acquisition module.
7. An audio - video control method with a wireless microphone intelligent tracking function according to claim 6, characterized in that: the spatial information acquisition module includes an audio information acquisition module and a video information acquisition module; the audio information acquisition module acquires audio information, and the audio information includes first audio information and second audio information; the video information acquisition module acquires video information, and the video information acquisition module includes a global human image information acquisition module and a local human image information acquisition module; the audio information acquisition module includes a digital signal conversion module, a vocal fluctuation frequency index calculation module, a turning digital signal marking module, and an audio information division module; the digital signal conversion module is used to convert audio information into digital signals, and the vocal fluctuation frequency index calculation module obtains the total time interval t0 between adjacent digital signals and the total information length p0 of the digital signals, and calculates the overall vocal fluctuation frequency index w = p0 / t0 of the digital signals; the turning digital signal marking module traverses the vocal fluctuation frequencies of the first digital signal and adjacent digital signals, and obtains the frequency fluctuation difference by taking the difference between the vocal fluctuation frequency of the first digital signal and the overall vocal fluctuation frequency index; sequentially obtains the vocal fluctuation frequencies of adjacent digital signals in the sound collection stage and marks the turning digital signals, where the turning digital signal is the digital signal corresponding to the quotient of the frequency fluctuation difference of the previous adjacent digital signal and the frequency fluctuation difference of the subsequent adjacent digital signal being negative, and the positive and negative of the frequency fluctuation differences of all digital signals after the turning digital signal are the same as the positive and negative of the frequency fluctuation difference corresponding to the turning digital signal; the audio information division module is used to divide the audio information before the turning digital signal into first audio information, and the audio information after the turning digital signal into second audio information.
8. An audio - video control method with a wireless microphone intelligent tracking function according to claim 6, characterized in that: the spatial information analysis module includes an audio information analysis module, a video information analysis module, and a person matching module; the audio information analysis module includes an audio attribute classification module, an average decibel difference ratio calculation module, and an audio attribute deviation index calculation module; the video information analysis module includes a global human image classification module, a person image ratio difference calculation module, and a global human image deviation index calculation module; the person matching module includes a deviation index similarity calculation module and a person - audio attribute correspondence module; the audio attribute analysis module classifies different audio attributes, and the average decibel difference ratio calculation module is used to record the decibel characteristics of each audio and denote different types of first audio attributes and their corresponding decibel characteristics as a set A, and respectively calculate the average decibel difference ratio of the decibel characteristics corresponding to the j - th type of first audio attribute in the set A changing with time; the audio attribute deviation index calculation module is used to calculate the overall deviation index of different types of first audio attributes in the set A. The global person image classification module is used to classify the person images in the global person image information acquisition module, and record the person image ratios corresponding to different person images as set B; the person image ratio difference calculation module is used to calculate the average person image ratio difference of different global person images in set B over time, and the global person image deviation index calculation module is used to calculate the overall deviation index of different global person images; The deviation index similarity calculation module is used to compare the numerical similarity between the global person image deviation index calculation module and the audio attribute deviation index calculation module. When it is greater than the similarity threshold, it indicates that the change in the person's decibel is related to the movement of the person's position in the global image; the person audio attribute correspondence module performs a one-to-one correspondence with a similarity greater than 99% between the average decibel difference ratio in set A and the average person image ratio difference in set B to obtain the audio attributes corresponding to different people in the global person image information.
9. A method for controlling audio and video with a wireless microphone intelligent tracking function according to claim 6, characterized in that: The position data acquisition module includes a person image ratio sorting module, a starting coordinate setting module, and a sector adaptation module; The person image ratio sorting module is used to take the person who first makes a sound in the second audio information as the starting person, and obtain the person ratios of all person images in the local person image information, and sort the person ratios from largest to smallest; the starting coordinate setting module sets the position coordinate corresponding to the person image with the smallest person ratio as the starting coordinate; The sector adaptation module is used to, when any person makes a sound during monitoring, based on the relationship between the person ratio of the person making a sound at this time in the local person image and the person ratio of the starting person, perform sector adaptation on the starting coordinate to obtain the position coordinate of the second person.
Citation Information
Patent Citations
Intelligent remote video conference system
CN111343411A
Conference portrait shooting method, interactive tablet, computer equipment and storage medium
CN112073613A