Conference control method and device based on AI vision and multi-device cooperation, equipment and storage medium
Through the combination of convolutional neural network and microphone array, the conference room environment is identified and the speaker location is determined, and the projector content is adjusted through speech recognition and semantic analysis, the problems of cumbersome equipment operation, poor coordination and insufficient recognition accuracy of AI vision technology in the existing technology are solved, and intelligent collaboration and automated control of multiple devices in the conference room is realized, and meeting efficiency and interactivity are improved.
Patent Information
- Application Number
- CN202510285298.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-20
AI Technical Summary
The existing conference control technology has problems such as cumbersome equipment operation, poor coordination, low degree of automation, and insufficient recognition accuracy and real-time performance of AI vision technology in complex scenarios.
The convolutional neural network is used to identify the conference room environment, determine the location information of the participants, and locate the sound source through the microphone array to determine the location of the speaker. Then, the camera is controlled toward the spokesperson, the speech content is identified through speech recognition and semantic analysis, and the projector content is adjusted according to the content.
It realizes intelligent collaboration and automated control of multiple devices in the conference room, improves the efficiency and interactivity of the meeting, and enhances the sense of participation and concentration of participants.
Smart Images

Figure CN120186293A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and particularly to a conference control method, device, equipment, and storage medium based on AI vision and multi-device collaboration. Background Art
[0002] In modern conference scenarios, with the increasing demand of enterprises for efficient collaboration and digital transformation, the intelligent upgrade of conference rooms has become an important trend. Traditional conference equipment and control methods are no longer sufficient to meet the complex and changing conference requirements. Therefore, there is an urgent need in the market for conference solutions that can achieve intelligent control to improve conference efficiency, optimize the user experience, and reduce equipment management costs.
[0003] Currently, in the industry, MAXHUB (a brand under Guangzhou Shiyuan Electronic Technology Co., Ltd.) has achieved comprehensive coverage of the conference space through a multi-modal multi-camera solution, combining a wide-angle lens and a telephoto lens. In addition, some enterprises are also exploring the use of AI vision technology for automatic identification and signing-in of participants to further optimize the conference process.
[0004] Although the existing practices have improved the intelligent level of conferences to a certain extent, there are still many deficiencies. For example, the device operation is cumbersome, the coordination between devices is poor, and the degree of automation of the conference process is low. In addition, the recognition accuracy and real-time performance of AI vision technology in complex scenarios still need to be improved. Especially in scenarios such as multi-person conferences and light changes, recognition errors or delays are likely to occur. Therefore, how to achieve intelligent collaboration and automatic control of multiple devices in a conference room has become an urgent problem to be solved.
[0005] The above content is only used to assist in understanding the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0006] The purpose of this application is to provide a conference control method, device, equipment, and storage medium based on AI vision and multi-device collaboration, aiming to solve the technical problem of how to achieve intelligent collaboration and automatic control of multiple devices in a conference room.
[0007] To achieve the above purpose, this application proposes a conference control method based on AI vision and multi-device collaboration, and the method includes:
[0008] Identifying the conference room environment through a convolutional neural network to determine the position information of the participants;
[0009] Performing sound source localization on the voice source of the microphone array according to the position information to determine the position of the speaker;
[0010] Controlling the camera to face the position of the speaker so that the display shows the speaker;
[0011] When the speaker is displayed on the display, the speech content of the speaker is recognized through speech recognition and semantic analysis, and the projector content is adjusted according to the speech content.
[0012] In one embodiment, the step of performing sound source localization on the speech source of the microphone array according to the position information to determine the speaker position includes:
[0013] Calculate the time difference of arrival between the speech signals captured by each microphone in the microphone array;
[0014] Determine the propagation direction of the speech signal according to the time difference of arrival;
[0015] Perform triangulation according to the time difference of arrival, the position information, the position of the microphone array, and the propagation direction to determine the speaker position.
[0016] In one embodiment, the step of performing triangulation according to the time difference of arrival, the position information, the position of the microphone array, and the propagation direction to determine the speaker position includes:
[0017] Calculate the propagation distance difference according to the time difference of arrival and the sound propagation speed;
[0018] Establish a speaker position equation according to the propagation distance difference and the position of the microphone array;
[0019] Numerically solve the speaker position equation by the least squares method to obtain the initial speaker position;
[0020] Adjust the initial speaker position according to the position information and the propagation direction to obtain the speaker position.
[0021] In one embodiment, after the step of, when the speaker is displayed on the display, recognizing the speech content of the speaker through speech recognition and semantic analysis, and adjusting the projector content according to the speech content, further includes:
[0022] Perform face recognition, pose recognition, and behavior analysis on the participants to obtain the emotional reactions of the participants to the meeting content;
[0023] Generate a feedback report according to the emotional reactions, and send the feedback report to a preset device;
[0024] Record the emotional reactions and the corresponding participants during the entire meeting process;
[0025] Generate a meeting effect report according to the emotional reactions and the corresponding participants.
[0026] In one embodiment, the step of, when the speaker is being displayed on the display, identifying the speech content of the speaker through speech recognition and semantic analysis and adjusting the projector content according to the speech content includes:
[0027] When the speaker is being displayed on the display, acquiring the voice signal captured by the microphone array;
[0028] Performing speech recognition on the voice signal through a deep learning language model to obtain a speech text;
[0029] Performing word segmentation, sentence segmentation, and part-of-speech tagging on the speech text to obtain the speech content of the speaker;
[0030] Performing semantic analysis on the speech content to obtain speech keywords;
[0031] Adjusting the projector content according to the speech keywords.
[0032] In one embodiment, after the step of, when the speaker is being displayed on the display, identifying the speech content of the speaker through speech recognition and semantic analysis and adjusting the projector content according to the speech content, the method further includes:
[0033] Performing text transcription on the speech content to obtain a transcription result;
[0034] Performing semantic analysis on the transcription result to obtain discussion topics and decision-making matters;
[0035] Annotating the corresponding speaker, the discussion topics, and the decision-making matters in the transcription result to form a meeting minutes.
[0036] In one embodiment, after the step of performing sound source localization on the voice source of the microphone array according to the position information to determine the speaker's position, the method further includes:
[0037] Adjusting the conference room lights according to the position information and the speaker's position;
[0038] When the adjustment of the conference room lights is completed, identifying the volume of the speaker to obtain an identification result;
[0039] Adjusting the volume of the microphone array and the sound system according to the identification result.
[0040] In addition, to achieve the above object, the present application further provides a conference control device based on AI vision and multi-device collaboration, and the device includes:
[0041] An environment recognition module, configured to identify the conference room environment through a convolutional neural network and determine the position information of the participants;
[0042] A positioning module, configured to perform sound source localization on the voice source of the microphone array according to the position information to determine the position of the speaker.
[0043] A camera control module, configured to control the camera to face the position of the speaker so that the speaker is displayed on the display.
[0044] An analysis module, configured to, when the speaker is displayed on the display, identify the speech content of the speaker through speech recognition and semantic analysis, and adjust the content of the projector according to the speech content.
[0045] In addition, to achieve the above object, the present application further provides a conference control device based on AI vision and multi-device collaboration. The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the conference control method based on AI vision and multi-device collaboration as described above.
[0046] In addition, to achieve the above object, the present application further provides a storage medium. The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the conference control method based on AI vision and multi-device collaboration as described above.
[0047] In addition, to achieve the above object, the present application further provides a computer program product. The computer program product includes a computer program. When the computer program is executed by a processor, it implements the steps of the conference control method based on AI vision and multi-device collaboration as described above.
[0048] One or more technical solutions proposed by the present application have at least the following technical effects:
[0049] The conference system first analyzes the images captured by the conference room cameras through a convolutional neural network to identify the environmental layout in the conference room and the location information of the participants, providing accurate data support for subsequent device control. Next, the system performs sound source localization on the voice signals captured by the microphone array to determine the specific location of the speaker. This process precisely calculates the direction and distance of the sound source, ensuring that the conference system can accurately identify the location of the speaker, providing a basis for the precise focusing of the camera and audio collection, thereby optimizing the audio capture effect. Subsequently, based on the speaker's location information, the system automatically adjusts the orientation of the camera to align it with the speaker and displays the speaker's image on the monitor in real time. This process is achieved through the pan-tilt control and autofocus functions of the camera, ensuring that the participants can clearly see the speaker's expressions and movements, enhancing the visual experience of the conference, and improving the interactivity and concentration of the conference. Finally, while the speaker's image is being displayed on the monitor, the system converts the speaker's voice signal into text through speech recognition technology and extracts the key information in the speech content using semantic analysis technology. Based on this information, the system dynamically adjusts the content of the projector, such as switching to relevant presentation pages or charts, thus realizing the real-time synchronization and intelligent display of the conference content. This process not only improves the efficiency of the conference but also enhances the participants' sense of participation and interactivity, realizing the intelligent collaboration and automated control of multiple devices in the conference room, providing strong support for the smooth progress of the conference. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0051] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0052] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the conference control method based on AI vision and multi-device collaboration of the present application;
[0053] Figure 2 It is a schematic flowchart provided for Embodiment 2 of the conference control method based on AI vision and multi-device collaboration of the present application;
[0054] Figure 3 It is a schematic diagram of the module structure of the conference control device based on AI vision and multi-device collaboration according to the embodiments of the present application;
[0055] Figure 4It is a schematic diagram of the device structure of the hardware operating environment involved in the conference control method based on AI vision and multi-device collaboration in the embodiments of the present application.
[0056] The implementation, functional features, and advantages of the present application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. Detailed implementation manners
[0057] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0058] To better understand the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0059] With the increasing demand of enterprises for efficient collaboration and digital transformation, intelligent upgrading has become an important trend in conference rooms. Traditional devices are difficult to meet the needs of modern conferences, prompting the market to seek more intelligent solutions to improve efficiency, optimize the experience, and reduce costs. Currently, companies such as MAXHUB achieve comprehensive space coverage through multi-modal camera technology, and some enterprises are exploring AI vision for automatic check-in to optimize processes. However, existing solutions still face problems such as complex device operations, poor coordination, low automation, and insufficient accuracy and real-time performance of AI recognition in complex scenarios.
[0060] The main solution of the embodiments of the present application is that the conference system uses a convolutional neural network to analyze camera images to identify the conference room layout and personnel positions, performs sound source localization through a microphone array to determine the position of the speaker, and accordingly adjusts the camera to face the speaker to optimize audio capture and visual experience. At the same time, the system converts speech into text and extracts key information through semantic analysis, dynamically adjusts the projection content such as switching presentation pages, realizes real-time synchronization and intelligent display of conference content, and enhances interactivity and concentration.
[0061] It should be noted that the execution subject of the embodiments of the present application can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a conference control system, etc. that can implement the above functions. Hereinafter, the conference control system will be taken as an example to illustrate this embodiment and the following embodiments.
[0062] Based on this, the embodiments of the present application provide a conference control method based on AI vision and multi-device collaboration, with reference to Figure 1 , Figure 1 It is a schematic flowchart of the first embodiment of the conference control method based on AI vision and multi-device collaboration of the present application.
[0063] In this embodiment, the conference control method based on AI vision and multi-device collaboration includes steps S10 to S40:
[0064] Step S10: Identify the meeting room environment through a convolutional neural network and determine the location information of the participants.
[0065] It should be noted that CNN (Convolutional Neural Network) is a deep learning model mainly used to process data with grid structures, such as images. Through structures such as convolutional layers, pooling layers, and fully connected layers, CNN extracts the features of the input data layer by layer, and can automatically learn the spatial hierarchical features in the image. It has a wide range of applications in the field of computer vision, including image classification, object detection, and face recognition. To train CNN for determining the location information of the participants, a large amount of image data labeled with location information needs to be prepared first. These images contain different postures and locations of the meeting room scene and the participants. Then, use these data to train CNN, and gradually extract the image features through convolutional layers and pooling layers to learn the appearance features of the participants and their locations in the image. During the training process, CNN will implicitly learn the location information, especially when using the zero-padding operation, which helps the network perceive boundary effects and extract location-related features. Finally, after multiple iterations of optimization, CNN can accurately identify the location information of the participants from the input image, providing support for subsequent meeting control and management.
[0066] The meeting room environment refers to the physical space inside the meeting room and its internal layout and equipment configuration, which includes the size, shape, lighting conditions of the meeting room, and various devices installed indoors (such as cameras, projectors, displays, sensors, etc.). These factors together constitute a complex visual scene that needs to be perceived and analyzed through AI vision technology. The participants refer to the people attending the meeting or event, such as the host, speaker, audience, etc.
[0067] The location information refers to the specific location of the participants inside the meeting room, usually represented by coordinates (such as two-dimensional coordinates or three-dimensional coordinates). The location information can be captured by a camera to obtain an image, and then analyzed and recognized using a convolutional neural network. For example, through image recognition technology, it is possible to determine the seat position of the participants in the meeting room, whether they are standing or moving, etc.
[0068] It is understandable that, first, the conference control system collects image data in real time through cameras in the conference room, and these images contain the layout of the conference room and the appearance information of the participants. Second, the collected images are input into a pre-trained convolutional neural network model, and using its powerful feature extraction ability, the environmental features in the conference room and the outlines of the participants are automatically recognized to determine the position information of each person. Finally, the recognized position information is combined with the layout of the conference room to generate the precise position coordinates of the participants.
[0069] Step S20: Perform sound source localization on the sound source of the voice of the microphone array according to the position information to determine the position of the speaker.
[0070] It should be noted that the microphone array refers to a system composed of multiple microphones arranged in a certain geometric pattern. These microphones work together to collect sound signals coming from different directions, and use signal processing technology to enhance the target sound and suppress noise. The layout forms of the microphone array are diverse, and common ones include linear arrays, circular arrays, and matrix arrays, etc. Different layouts are suitable for different application scenarios.
[0071] The sound source of the voice refers to the emission point of the sound signal, that is, the position where the sound is generated. In the conference room scenario, the sound source of the voice is usually the voices of the participants.
[0072] Sound source localization is to measure the sound signal using multiple microphones, and by analyzing the time delay (TDOA) of the sound signal arriving at different microphones or other features, calculate the position information of the sound source relative to the microphone array, including azimuth angle, elevation angle, and distance, etc. Sound source localization methods include the TDOA method based on arrival time difference estimation, the method based on beamforming, and the method based on deep learning.
[0073] The position of the speaker refers to the specific position of the sound source in physical space, usually represented in the form of coordinates. Through sound source localization technology, the azimuth and distance of the speaker relative to the microphone array can be determined, so as to achieve precise tracking of the position of the speaker.
[0074] It is understandable that, first, the conference control system uses the microphone array to collect voice signals. The microphone array captures the sound signal and its propagation time difference through the collaborative work of multiple microphones. Second, the system calculates the time difference of the sound source arriving at different microphones based on the time difference of arrival algorithm, combines the geometric layout of the microphones and the speed of sound, and deduces the position of the sound source through geometric relationships. Finally, the conference control system combines the sound source localization result with the layout information of the conference room to determine the exact position of the speaker, so as to achieve precise tracking and positioning of the speaker and enhance the intelligent experience of the conference.
[0075] As an example, after the step of performing sound source localization on the voice source of the microphone array according to the position information to determine the speaker's position, the following steps are further included: adjusting the conference room lighting according to the position information and the speaker's position; when the adjustment of the conference room lighting is completed, identifying the volume of the speaker to obtain an identification result; and adjusting the volumes of the microphone array and the sound system according to the identification result.
[0076] The conference room lighting refers to the lighting system in the conference room, including the brightness, color temperature, layout, and control method of the lamps. Through the intelligent lighting system, the brightness and color temperature of the lighting can be automatically or manually adjusted according to different scene requirements of the meeting to create a suitable visual environment.
[0077] The identification result refers to the volume data obtained by analyzing the speaker's volume through the sound source localization technology. This data reflects the sound intensity of the speaker and can be used to further adjust the volumes of the microphone and the sound system.
[0078] First, the conference control system automatically adjusts the conference room lighting according to the determined position information of the participants and the speaker's position. Through the intelligent lighting system, the lighting is focused on the area where the speaker is located, and at the same time, the brightness and color temperature are adjusted to highlight the speaker and create a suitable conference atmosphere, avoiding affecting the conference effect due to too strong or too weak lighting. Second, after the adjustment of the conference room lighting is completed, the system starts the volume identification module to monitor and analyze the speaker's voice signal in real time. The sound is collected through the microphone array and its volume is calculated to obtain the volume identification result. Finally, according to the volume identification result, the conference control system dynamically adjusts the sensitivity of the microphone array and the output volume of the sound system to ensure that the speaker's voice can be clearly and evenly transmitted to every corner of the conference room, while avoiding the voice being too strong or too weak, thereby realizing the intelligent optimization of the conference environment and improving the fluency and sense of participation of the conference.
[0079] Step S30: Control the camera to face the speaker's position so that the display shows the speaker.
[0080] It should be noted that the camera refers to a video capture device installed in the conference room, which is used to collect image and video information in the conference scene in real time. It usually has an adjustable viewing angle and focal length and can be aimed at a specific area or person, such as the speaker, according to needs. In the intelligent conference system, the camera may support functions such as autofocus, zoom, and pan-tilt control to achieve remote control and precise positioning.
[0081] The display refers to a device for displaying video signals, such as a TV or a display screen in the conference room. It receives signals from the camera or other video sources and converts them into visual images or video content for the participants to watch. The display is used to show the image of the speaker or other multimedia content in the conference.
[0082] It can be understood that, firstly, based on the determined speaker position information, the conference control system sends instructions through the pan-tilt control module to make the camera rotate quickly and accurately aim at the speaker, ensuring that the camera's view covers the whole body or upper body of the speaker, so as to obtain a clear picture. Secondly, after the camera captures the image of the speaker, it transmits the video signal to the connected display in real time. After receiving the signal, the display immediately shows the image of the speaker, ensuring that the participants can clearly see the speaker. Finally, in this way, the conference control system realizes the coordinated control of the camera and the display, not only improving the visual effect of the conference, but also enhancing the concentration and sense of participation of the participants, making the conference more efficient and smooth.
[0083] Step S40, when the speaker is displayed on the said display, identify the speech content of the said speaker through speech recognition and semantic analysis, and adjust the projector content according to the said speech content.
[0084] It should be noted that semantic analysis refers to the process of processing natural language text to understand its meaning. It extracts the core meaning and intention of the text by analyzing the words, grammatical structures and context relationships in the text, so as to achieve a deep understanding of the language.
[0085] Projector content refers to the image or video information displayed by the projector, usually including presentation documents, charts, pictures, videos or other multimedia content. In a conference scenario, projector content is an important carrier for conference information transmission, used to assist the speaker's explanation and help the participants better understand and absorb information.
[0086] It can be understood that, firstly, when the speaker's image is already displayed on the display, the conference control system activates the speech recognition module to capture the speaker's speech signal in real time and convert it into text data. Subsequently, the system performs semantic analysis on the generated text to extract the speech content. Finally, according to the speech content, the content of the projector is adjusted, switching to the page related to the speech content or displaying relevant charts, pictures, etc., so as to realize the dynamic synchronization of the projector content with the speaker's explanation, enhancing the interactivity and information transmission efficiency of the conference.
[0087] As an example, the steps of, when the speaker is displayed on the display, identifying the speech content of the speaker through speech recognition and semantic analysis and adjusting the projector content according to the speech content include: when the speaker is displayed on the display, acquiring a speech signal captured by a microphone array; performing speech recognition on the speech signal through a deep learning language model to obtain a speech text; performing word segmentation, sentence segmentation, and part-of-speech tagging on the speech text to obtain the speech content of the speaker; performing semantic analysis on the speech content to obtain speech keywords; and adjusting the projector content according to the speech keywords.
[0088] A speech signal refers to the voice data of a speaker captured by a microphone array, which usually exists in the form of an audio signal and contains the speech information of the speaker.
[0089] A deep learning language model refers to a language processing model constructed based on deep learning techniques, such as models with the Transformer architecture (such as BERT, GPT series). These models can efficiently recognize and understand speech signals and generate accurate text outputs.
[0090] A speech text refers to the text content obtained by converting a speech signal through speech recognition technology. It is the textual representation of the speech signal and is used for subsequent text processing and analysis.
[0091] Word segmentation refers to the process of splitting continuous text in a speech text into independent lexical units. For example, splitting "I love natural language processing" into "I / love / natural language processing" to facilitate better semantic analysis.
[0092] Sentence segmentation refers to splitting a speech text into independent sentences according to semantic and syntactic rules. For example, splitting a continuous speech text into multiple complete sentences for sentence-by-sentence analysis and processing.
[0093] Part-of-speech tagging refers to classifying the words after word segmentation into grammatical categories and tagging the part of speech of each word (such as noun, verb, adjective, etc.) to help understand the structure and semantics of the sentence.
[0094] The speech content refers to the speech text after word segmentation, sentence segmentation, and part-of-speech tagging processing, which contains the complete semantic information of the speaker.
[0095] Speech keywords refer to the core words or phrases extracted from the speech content through semantic analysis. These keywords can reflect the core theme and key information of the speech and are used for subsequent adjustment of the projector content.
[0096] First, when the speaker is displayed on the monitor, the conference control system captures the speaker's voice signal through the microphone array to provide high-quality audio data for subsequent processing. Then, the system uses a deep learning language model to perform speech recognition on the captured voice signal and converts the voice signal into accurate speech text. Next, the system performs natural language processing on the generated speech text, including word segmentation, sentence segmentation, and part-of-speech tagging, splitting the continuous text into independent words and sentences and tagging the part of speech of each word, thus obtaining the complete speech content, which provides a basis for semantic analysis. Subsequently, the system performs semantic analysis on the speech content to extract speech keywords, which can reflect the core theme and key content of the speech. Finally, based on the extracted speech keywords, the system adjusts the content of the projector, switches to the page related to the speech content or displays relevant materials, realizing the dynamic synchronization between the projector content and the speaker's explanation, and improving the interactivity and information transmission efficiency of the conference.
[0097] As an example, after the steps of recognizing the speaker's speech content through speech recognition and semantic analysis and adjusting the projector content according to the speech content when the speaker is displayed on the monitor, the method further includes: performing face recognition, posture recognition, and behavior analysis on the participants to obtain the emotional reactions of the participants to the conference content; generating a feedback report according to the emotional reactions and sending the feedback report to a preset device; recording the emotional reactions and the corresponding participants during the entire conference process; and generating a conference effect report according to the emotional reactions and the corresponding participants.
[0098] Face recognition refers to using computer vision technology to analyze face images, and realizing identity recognition or verification by detecting face features (such as the positions of eyes, nose, and mouth and their relative distances) and extracting feature vectors.
[0099] Posture recognition is to analyze human postures and movements through computer vision technology, extract human key points (such as joints, head, etc.) and skeleton information, so as to recognize human postures and movements.
[0100] Behavior analysis refers to analyzing the behavior patterns of individuals or groups through data-driven methods to understand their behavior characteristics and potential intentions. In a conference scenario, behavior analysis can be used to evaluate the participation and reactions of the participants.
[0101] Emotional reaction refers to the inner experience and emotional response generated by an individual when facing external stimuli, usually manifested as changes in facial expressions, body movements, or voice intonations. In a conference, emotional reactions can reflect the acceptance degree and emotional state of the participants towards the conference content.
[0102] A feedback report refers to a summary document generated during a meeting based on real-time collected emotional reactions and behavioral data of the participants. It is used to provide immediate feedback to the meeting host to help them understand the acceptance level and emotional state of the participants towards the meeting content. Such a report usually includes information such as the facial expressions, postures, and interaction frequencies of the participants, which can help the host timely adjust the meeting progress, content, or method to better meet the needs of the participants.
[0103] A preset device refers to the terminal device pre-set by the meeting system to receive the feedback report, such as the tablet computer, laptop, or meeting console used by the host. These devices can receive the feedback report in real time and display the relevant information through a visual interface for the host to make quick decisions. The selection and configuration of the preset device are aimed at ensuring that the feedback information can be conveyed to the leader of the meeting timely and accurately.
[0104] The meeting effectiveness report is a summary document generated based on the detailed data collected during the entire meeting process and is used to evaluate the overall effectiveness of the meeting. It not only includes the analysis results of the emotional reactions and behaviors of the participants but may also cover aspects such as the achievement of the meeting objectives, the depth and breadth of the discussions, and the time management efficiency. It is usually generated after the meeting to summarize the successful experiences and deficiencies of the meeting and provide improvement suggestions for future meetings.
[0105] First, the meeting system collects the image and video data of the participants in real time and analyzes their facial expressions, body postures, and behavioral actions through computer vision algorithms to accurately judge the emotional reactions of the participants towards the meeting content, such as concentration, confusion, or positive feedback, providing a basis for subsequent adjustments. Second, the system generates a feedback report based on these emotional reactions, presenting it in an intuitive chart or text form and pushing it to the preset host device in real time to help the host immediately understand the emotions and participation status of the participants, so as to timely adjust the meeting rhythm or content and enhance the interactivity and effectiveness of the meeting. At the same time, the system will record the emotional reactions of the participants and their corresponding identity information throughout the process to retain data for subsequent detailed analysis. Finally, based on the recorded data, the system generates a meeting effectiveness report to comprehensively evaluate the overall effectiveness of the meeting, including the participation degree of the participants, the acceptance level of the meeting content, and potential improvement directions, providing data support and improvement suggestions for future meeting optimization.
[0106] As an example, after the step of, when the speaker is displayed on the display, recognizing the speech content of the speaker through speech recognition and semantic analysis and adjusting the projector content according to the speech content, the method further includes: transcribing the speech content into text to obtain a transcription result; performing semantic analysis on the transcription result to obtain discussion topics and decision items; and marking the corresponding speaker, the discussion topics, and the decision items in the transcription result to form a meeting minutes.
[0107] The transcription result refers to the text record obtained by converting the speech content of the speaker into text form through speech recognition technology, which completely reflects the speech content of the speaker in the meeting.
[0108] The discussion topic refers to the theme or issue proposed and centrally discussed in the meeting. These topics are usually the core content of the meeting, reflecting the focus and goal of the meeting, such as project progress, market strategy, or technical solution, etc.
[0109] The decision item refers to the conclusion, decision, or action plan formed during the meeting discussion. These items are usually responses to the discussion topics, clarifying the next action direction or specific decision result, such as approving a certain plan, assigning tasks, or setting time nodes, etc.
[0110] First, the meeting system converts the speech signal of the speaker into text through speech recognition technology to generate a transcription result. The speech recognition process includes speech signal acquisition, preprocessing, feature extraction, and pattern recognition. The system uses a microphone to collect the speech signal. After noise removal and feature extraction (such as Mel Frequency Cepstral Coefficients, MFCC), it performs pattern recognition through deep learning models (such as convolutional neural network and recurrent neural network), and finally converts the speech into text. Secondly, the system performs semantic analysis on the transcription result to extract discussion topics and decision items. Semantic analysis uses deep learning methods (such as Transformer model) or knowledge graph-based technologies to identify key information in the text, including entity recognition, relation extraction, and semantic role annotation, which can help the system understand the semantic structure of the text, so as to accurately extract the core topics and decision content in the meeting. Finally, the system marks the corresponding speaker, discussion topics, and decision items in the transcription result to form a clear meeting minutes. This process is achieved through natural language processing technology, integrating and formatting the results of speech recognition and semantic analysis to ensure that the meeting minutes accurately reflect the meeting content and are convenient for subsequent reference and execution.
[0111] This embodiment provides a meeting control method based on AI vision and multi-device collaboration. First, the meeting system analyzes the images collected by the conference room camera through a convolutional neural network to identify the environmental layout in the meeting room and the position information of the participants, providing accurate data support for subsequent device control. Then, the system performs sound source localization on the voice signals captured by the microphone array to determine the specific position of the speaker. This process accurately calculates the direction and distance of the sound source to ensure that the meeting system can accurately identify the position of the speaker, providing a basis for the accurate focusing and audio collection of the camera, thereby optimizing the audio capture effect. Subsequently, based on the position information of the speaker, the system automatically adjusts the orientation of the camera to align it with the speaker and displays the speaker's image on the monitor in real time. This process is achieved through the pan-tilt control and autofocus functions of the camera, ensuring that the participants can clearly see the speaker's expressions and movements, enhancing the visual experience of the meeting, and improving the interactivity and concentration of the meeting. Finally, while the speaker's image is being displayed on the monitor, the system converts the speaker's voice signal into text through speech recognition technology and extracts the key information in the speech content using semantic analysis technology. Based on this information, the system dynamically adjusts the content of the projector, such as switching to relevant presentation pages or charts, thereby achieving real-time synchronization and intelligent display of the meeting content. This process not only improves the efficiency of the meeting but also enhances the participants' sense of participation and interactivity, realizing the intelligent collaboration and automated control of multiple devices in the meeting room and providing strong support for the smooth progress of the meeting.
[0112] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 , Figure 2 is a schematic flowchart of the second embodiment of the meeting control method based on AI vision and multi-device collaboration of the present application. The step S20 of the meeting control method based on AI vision and multi-device collaboration includes steps S21 to S23:
[0113] Step S21, calculate the time difference of arrival between the voice signals captured by each microphone in the microphone array.
[0114] It should be noted that the time difference of arrival refers to the time difference required for sound waves to propagate from the sound source to different microphones in the microphone array. Since the speed of sound waves propagating in the air is limited, when the sound source emits sound, the sound will reach each microphone in the array at different times. This time difference is caused by the distance difference between the sound source and each microphone. By calculating these time differences of arrival, the system can further infer the direction and position of the sound source, thereby realizing the function of sound source localization.
[0115] It is understandable that, first, each microphone in the microphone array simultaneously captures the voice signal emitted by the sound source. Since the distances from the sound source to each microphone are different, the arrival times of the sound waves at each microphone will also vary. The system precisely measures the moments when each microphone receives the voice signal and calculates the time difference in the arrival of the voice signal between any two microphones, that is, the time difference of arrival. This process typically uses the Generalized Cross-Correlation (GCC-PHAT) algorithm, which estimates the time delays of the sound source arriving at different microphones by analyzing the correlation between signals. Assume that the signals received by two microphones are x(t) and y(t) respectively, and their cross-correlation function is:
[0116] R xy (τ)=∫x(t)y(t + τ)dt
[0117] Through the GCC-PHAT algorithm, the time difference τ of the sound source arriving at the two microphones can be obtained, that is:
[0118]
[0119] Step S22, determine the propagation direction of the voice signal according to the time difference of arrival.
[0120] It should be noted that the propagation direction refers to the path direction of the sound wave from the sound source to the microphone array. In sound source localization, by calculating the time differences of arrival between the voice signals captured by each microphone in the microphone array, the propagation direction of the sound wave can be inferred.
[0121] It is understandable that the system constructs a hyperbola equation set based on the time difference of arrival and the known geometric layout of the microphone array, and solves these equations by the least squares method or other mathematical methods to determine the propagation direction of the sound source. Based on the time difference of arrival τ and the speed of sound v, the propagation direction of the sound source can be calculated. For two microphones, the arrival angle θ of the sound source can be calculated by the following formula:
[0122]
[0123] where d is the distance between the two microphones.
[0124] Step S23, perform triangulation according to the time difference of arrival, the position information, the position of the microphone array, and the propagation direction to determine the speaker position.
[0125] It should be noted that triangulation refers to determining the position of an unknown point by using the known geometric relationships and measurement data and the properties of triangles.
[0126] It can be understood that in sound source localization, triangulation is based on position information such as the time difference of arrival and the geometric layout of the microphone array. By calculating the distance differences between the sound source and each microphone, and combining the geometric relationships of triangles, the specific position of the sound source can be deduced.
[0127] As an example, the step of triangulating according to the time difference of arrival, the position information, the positions of the microphone array, and the propagation direction to determine the speaker position includes: calculating the propagation distance difference according to the time difference of arrival and the sound propagation speed; establishing a speaker position equation according to the propagation distance difference and the positions of the microphone array; numerically solving the speaker position equation by the least squares method to obtain an initial speaker position; and adjusting the initial speaker position according to the position information and the propagation direction to obtain the speaker position.
[0128] The propagation distance difference refers to the distance difference between the sound waves emitted by the sound source when reaching different microphones in the microphone array. Through the known time difference of arrival and the sound propagation speed, the distance differences from the sound source to each microphone can be calculated, and this distance difference is a key parameter used in triangulation to determine the position of the sound source.
[0129] The speaker position equation refers to a mathematical equation established through geometric relationships based on the propagation distance difference and the known positions of the microphone array. These equations describe the relationship between the sound source position and the microphone positions, usually a system of equations containing multiple variables. By solving these equations, the coordinates of the sound source, that is, the speaker position, can be obtained.
[0130] The initial speaker position refers to the rough position obtained by numerically solving the speaker position equation by a numerical method (such as the least squares method). This position is an approximate value used for subsequent precise adjustment. Due to the existence of noise and errors in actual measurements, the initial position may not be precise enough, so further optimization is required.
[0131] First, calculate the propagation distance differences between the sound source and each microphone through the time difference of arrival and the speed of sound. For example, assume the speed of sound is 340 m / s and the time difference of arrival between microphone A and microphone B is 1 ms, then the distance difference between the sound source and these two microphones is 0.34 m. Secondly, use these propagation distance differences and the known positions of each microphone in the microphone array to establish a system of geometric equations. Assume microphone A and microphone B are located at coordinates (0,0) and (1,0) respectively, and the distances from the sound source to microphone A and microphone B are d A and d B , then the following system of equations can be established:
[0132]
[0133] where x and y are the coordinates of the sound source.
[0134] Then, numerically solve the above system of equations by the least squares method to obtain the initial speaker position. Transform the above equations into matrix form and solve them by the least squares method:
[0135] x = (M T M) -1 M T ·b
[0136] where M is the microphone position matrix, b is the propagation distance difference vector, and x is the coordinate of the sound source.
[0137] Suppose the solution result is the coordinate (0.5, 0.3), which is a rough position estimate and may have errors. Finally, adjust the initial speaker position according to the known positions of the participants and the propagation direction of the sound source. For example, assume that through image analysis of the camera, it is known that the participants are roughly in the left area of the room, while the initial position (0.5, 0.3) is in the center of the room. Combining the propagation direction information (assuming the sound source mainly faces microphone A), it can be judged that the initial position may be too far to the right. Therefore, adjust the initial position to the left, for example, to (0.3, 0.3). In this way, combine multi-source information to correct the initial position and finally obtain a more accurate speaker position. For example, assume the initial position is (x0, y0) and the propagation direction is θ, then the position can be adjusted by the following formula:
[0138] x final = x0 + Δx·cos(θ)
[0139] y final = y 0 + Δy·sin(θ)
[0140] where Δx and Δy are the adjustment amounts based on the position information.
[0141] In this embodiment, the time difference of arrival between the voice signals captured by each microphone in the microphone array is first calculated. By accurately measuring the time difference of the sound wave arriving at different microphones, key data is provided for subsequent positioning. Then, the propagation direction of the voice signal is determined according to the time difference of arrival. By analyzing the time sequence of the sound wave arriving at different microphones, the general direction of the sound source is inferred, effectively narrowing the possible position range of the sound source and providing a direction guidance for subsequent precise positioning. Finally, triangulation is performed by combining the time difference of arrival, the position information of the participants, the known position of the microphone array, and the propagation direction. By constructing and solving a geometric equation set, the precise position of the speaker is determined. This process utilizes multi-source information, further improving the positioning accuracy, enabling fast and accurate positioning of the speaker, providing support for functions such as automatic camera focusing and audio acquisition optimization in the conference system, and significantly enhancing the intelligent level and user experience of the conference.
[0142] It should be noted that the above examples are only for understanding this application and do not constitute a limitation to the conference control method based on AI vision and multi-device collaboration of this application. Any simple transformation in more forms based on this technical concept is within the protection scope of this application.
[0143] This application also provides a conference control device based on AI vision and multi-device collaboration. Please refer to Figure 3 , the conference control device based on AI vision and multi-device collaboration includes:
[0144] An environment recognition module 10, configured to recognize the conference room environment through a convolutional neural network and determine the position information of the participants;
[0145] A positioning module 20, configured to perform sound source positioning on the voice source of the microphone array according to the position information and determine the position of the speaker;
[0146] A camera control module 30, configured to control the camera to face the position of the speaker so that the speaker is displayed on the display;
[0147] An analysis module 40, configured to, when the speaker is displayed on the display, recognize the speech content of the speaker through speech recognition and semantic analysis and adjust the content of the projector according to the speech content.
[0148] In one embodiment, the positioning module 20 is further configured to calculate the time difference of arrival between the voice signals captured by each microphone in the microphone array; determine the propagation direction of the voice signal according to the time difference of arrival; perform triangulation according to the time difference of arrival, the position information, the position of the microphone array, and the propagation direction to determine the position of the speaker.
[0149] In one embodiment, the positioning module 20 is further configured to calculate a propagation distance difference according to the time difference of arrival and the sound propagation speed; establish a speaker position equation based on the propagation distance difference and the position of the microphone array; numerically solve the speaker position equation by the least squares method to obtain an initial speaker position; and adjust the initial speaker position according to the position information and the propagation direction to obtain the speaker position.
[0150] In one embodiment, the analysis module 40 is further configured to perform face recognition, pose recognition, and behavior analysis on the participants to obtain the emotional reactions of the participants to the meeting content; generate a feedback report according to the emotional reactions and send the feedback report to a preset device; record the emotional reactions and the corresponding participants during the entire meeting process; and generate a meeting effect report according to the emotional reactions and the corresponding participants.
[0151] In one embodiment, when the speaker is displayed on the display, the analysis module 40 is further configured to acquire a voice signal captured by the microphone array; perform speech recognition on the voice signal through a deep learning language model to obtain a speech text; perform word segmentation, sentence segmentation, and part-of-speech tagging on the speech text to obtain the speech content of the speaker; perform semantic analysis on the speech content to obtain speech keywords; and adjust the content of the projector according to the speech keywords.
[0152] In one embodiment, the analysis module 40 is further configured to transcribe the speech content to obtain a transcription result; perform semantic analysis on the transcription result to obtain discussion topics and decision-making matters; and mark the corresponding speaker, the discussion topics, and the decision-making matters in the transcription result to form a meeting minutes.
[0153] In one embodiment, the positioning module 20 is further configured to adjust the conference room lights according to the position information and the speaker position; recognize the volume of the speaker when the adjustment of the conference room lights is completed to obtain a recognition result; and adjust the volumes of the microphone array and the audio according to the recognition result.
[0154] The conference control device based on AI vision and multi-device collaboration provided by the present application adopts the conference control method based on AI vision and multi-device collaboration in the above embodiments, and can solve the technical problem of how to achieve intelligent collaboration and automatic control of multiple devices in a conference room. Compared with the prior art, the beneficial effects of the conference control device based on AI vision and multi-device collaboration provided by the present application are the same as those of the conference control method based on AI vision and multi-device collaboration provided by the above embodiments, and other technical features in the conference control device based on AI vision and multi-device collaboration are the same as the features disclosed in the above embodiment method, which will not be elaborated here.
[0155] The present application provides a conference control device based on AI vision and multi-device collaboration. The conference control device based on AI vision and multi-device collaboration includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the conference control method based on AI vision and multi-device collaboration in Embodiment 1 above.
[0156] Refer to the following Figure 4 , which shows a schematic structural diagram of a conference control device suitable for implementing the embodiments of the present application based on AI vision and multi-device collaboration. The conference control device based on AI vision and multi-device collaboration in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions, tablet computers), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The shown conference control device based on AI vision and multi-device collaboration is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0157] As Figure 4As shown, the conference control device based on AI vision and multi-device collaboration may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the program stored in the ROM (Read Only Memory) 1002 or the program loaded from the storage device 1003 into the RAM (Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the conference control device based on AI vision and multi-device collaboration are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, an LCD (Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the conference control device based on AI vision and multi-device collaboration to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a conference control device based on AI vision and multi-device collaboration with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.
[0158] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0159] The conference control device based on AI vision and multi-device collaboration provided by this application adopts the conference control method based on AI vision and multi-device collaboration in the above embodiments, and can solve the technical problem of how to achieve intelligent collaboration and automated control of multiple devices in a conference room. Compared with the prior art, the beneficial effects of the conference control device based on AI vision and multi-device collaboration provided by this application are the same as those of the conference control method based on AI vision and multi-device collaboration provided by the above embodiments, and other technical features in the conference control device based on AI vision and multi-device collaboration are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.
[0160] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0161] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0162] This application provides a computer-readable storage medium with computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the conference control method based on AI vision and multi-device collaboration in the above embodiments.
[0163] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory or flash memory), optical fibers, CD-ROM (Compact Disk - Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0164] The above computer-readable storage medium can be included in a conference control device based on AI vision and multi-device collaboration; it can also exist separately without being assembled into a conference control device based on AI vision and multi-device collaboration.
[0165] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by a conference control device based on AI vision and multi-device collaboration, the conference control device based on AI vision and multi-device collaboration is enabled to: recognize the conference room environment through a convolutional neural network and determine the position information of the participants; perform sound source localization on the voice source of the microphone array according to the position information to determine the speaker's position; control the camera to face the speaker's position so that the display shows the speaker; when the speaker is shown on the display, recognize the speech content of the speaker through speech recognition and semantic analysis, and adjust the projector content according to the speech content.
[0166] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a LAN (Local Area Network) or a WAN (Wide Area Network), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0167] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0168] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.
[0169] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned conference control method based on AI vision and multi-device collaboration, and can solve the technical problem of how to achieve intelligent collaboration and automated control of multiple devices in a conference room. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the conference control method based on AI vision and multi-device collaboration provided by the above embodiments, and will not be elaborated here.
[0170] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the above-mentioned conference control method based on AI vision and multi-device collaboration.
[0171] The computer program product provided by the present application can solve the technical problem of how to achieve intelligent collaboration and automated control of multiple devices in a conference room. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the above-mentioned conference control method based on AI vision and multi-device collaboration, and will not be elaborated here.
[0172] The above are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A conference control method based on AI vision and multi-device collaboration, characterized in that: The method comprises: Use convolutional neural networks to identify the conference room environment and determine the location information of the participants; Performing sound source positioning on the voice source of the microphone array according to the position information to determine the speaker's position; Controlling the camera to face the speaker's position so that the display shows the speaker; When the display shows the speaker, the speaker's speech content is identified through voice recognition and semantic analysis, and the projector content is adjusted according to the speech content.
2. The method according to claim 1, characterized in that The step of locating the voice source of the microphone array according to the position information to determine the speaker's position includes: Calculate the arrival time difference between the speech signals captured by each microphone in the microphone array; Determining a propagation direction of the speech signal according to the arrival time difference; The speaker position is determined by performing triangulation based on the arrival time difference, the position information, the position of the microphone array, and the propagation direction.
3. The method according to claim 2, characterized in that The step of determining the speaker's position by triangulating the arrival time difference, the position information, the position of the microphone array, and the propagation direction comprises: Calculate the propagation distance difference according to the arrival time difference and the sound propagation speed; Establishing a speaker position equation according to the propagation distance difference and the position of the microphone array; Numerically solving the speaker position equation by the least square method to obtain an initial speaker position; The initial speaker position is adjusted according to the position information and the propagation direction to obtain the speaker position.
4. The method according to claim 1, characterized in that After the step of identifying the speaker's speech content through voice recognition and semantic analysis and adjusting the projector content according to the speech content when the display displays the speaker, the method further includes: Performing facial recognition, gesture recognition, and behavior analysis on the participants to obtain the participants' emotional responses to the meeting content; Generating a feedback report according to the emotional response, and sending the feedback report to a preset device; Record the emotional reactions and corresponding participants throughout the meeting; A meeting effectiveness report is generated based on the emotional reactions and the corresponding meeting participants.
5. The method according to claim 1, characterized in that The step of identifying the speaker's speech content through speech recognition and semantic analysis, and adjusting the projector content according to the speech content when the display displays the speaker includes: When the display shows the speaker, acquiring a speech signal captured by a microphone array; Performing speech recognition on the speech signal through a deep learning language model to obtain speech text; Perform word segmentation, sentence segmentation and part-of-speech tagging on the speech text to obtain the speech content of the speaker; Performing semantic analysis on the speech content to obtain speech keywords; The projector content is adjusted according to the speech keywords.
6. The method according to claim 1, characterized in that After the step of identifying the speaker's speech content through voice recognition and semantic analysis and adjusting the projector content according to the speech content when the display displays the speaker, the method further includes: Transcribing the speech content to obtain a transcription result; Performing semantic analysis on the transcription results to obtain discussion topics and decision-making matters; The corresponding speakers, the discussion topics and the decision-making items are marked in the transcription results to form meeting minutes.
7. The method according to any one of claims 1 to 6, characterized in that After the step of locating the voice source of the microphone array according to the position information and determining the speaker's position, the method further includes: adjusting the conference room lighting according to the location information and the speaker's position; When the lighting of the conference room is adjusted, the speaker's volume is identified to obtain an identification result; The volume of the microphone array and the speaker is adjusted according to the recognition result.
8. A conference control device based on AI vision and multi-device collaboration, characterized in that: The device comprises: The environment recognition module is used to identify the conference room environment through a convolutional neural network and determine the location information of the participants; A positioning module, used to locate the sound source of the voice of the microphone array according to the position information, and determine the position of the speaker; A camera control module, used for controlling the camera to face the speaker position so that the display shows the speaker; The analysis module is used to identify the speech content of the speaker through voice recognition and semantic analysis when the display displays the speaker, and adjust the projector content according to the speech content.
9. A conference control device based on AI vision and multi-device collaboration, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the conference control method based on AI vision and multi-device collaboration as described in any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the conference control method based on AI vision and multi-device collaboration as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Target conference record voice processing method, device and equipment
CN121459821A