Split desktop conference terminal all-in-one machine

By designing a split desktop conference terminal all-in-one machine, it integrates a rotary conference host base, a ring array microphone and a voice recognition system, which solves the problems of poor interaction, poor sound pickup and complex wiring of existing equipment, and achieves an efficient and convenient meeting system experience and cost-reducing effect.

CN120050557APending Publication Date: 2025-05-27GUANGZHOU LIANGCHENG YUNQI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510203798.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing conference room desktop conference terminal equipment has problems such as poor interactive experience, poor sound pickup effect, complex wiring of equipment, easy loss of accessories, and high cost of use.

Method used

A split desktop conference terminal all-in-one machine is designed, including a base, a meeting host base, 8 microphones with a ring array, a voice recognition system, a touch tablet stand with magnetic charging and self-developed application software based on Android system. The device is connected through a rotating mechanism, has a rotating conference host base, supports 360-degree rotation, integrates a high-fidelity speaker and an 8-array microphone system, and has voice recognition and directional sound pickup functions.

Benefits of technology

It enables users to have a complete intelligent conference system experience without additional purchase of equipment, improves sound pickup and interactive experience, simplifies device wiring, reduces accessories loss and usage costs, and has a high voice recognition rate in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050557A_ABST
    Figure CN120050557A_ABST
Patent Text Reader

Abstract

The invention discloses a split type desktop conference terminal all-in-one machine, and belongs to the technical field of intelligent software and hardware equipment, the split type desktop conference terminal all-in-one machine comprises a base, a conference host base is arranged above the base, the base is connected with the conference host base through a rotating mechanism, eight microphones in an annular array are installed on the side face of the conference host base, and the rotating mechanism is connected with the rotating mechanism. The eight microphones are installed on the conference host base and used for picking up voice of a speaker in a conference, a voice recognition system used for controlling the eight microphones to conduct voice recognition is installed on the conference host base, and a touch panel support with a magnetic attraction charging function is arranged on the top of the conference host base and used for containing a touch conference control panel and charging the touch conference control panel. Compared with a traditional split-type scheme of a desktop microphone and a conference terminal, the device integrates the functions of the conference host and the control screen on the basis of the capability of the desktop microphone, and a user can have complete intelligent conference system experience without purchasing equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent software and hardware devices, and particularly relates to a split desktop conference terminal all-in-one machine. Background Art

[0002] Existing desktop conference terminals in meeting rooms generally have several forms:

[0003] Desktop omnidirectional microphones. Such devices generally integrate multiple microphones to pick up voices in the meeting room and generally exist as peripherals without an operating system and cannot run conference software.

[0004] Conference terminal hosts. Such conference terminal hosts generally can run conference software based on Windows, Android or other operating systems, but they do not have audio and video capabilities themselves and need to connect to peripherals to achieve audio and video collection.

[0005] Conference large screens. Such devices integrate a screen, a camera and an array microphone and can meet the audio and video requirements of the meeting room scenario. Generally, such devices also have an operating system and can directly run conference software.

[0006] Multimedia conference terminals. Such terminals generally integrate a camera and an array microphone and are placed on the desktop or above the TV. Some of such devices integrate an operating system and can send video signals to display devices such as TVs and projectors through HDMI to enable them to have audio and video and conference software running capabilities. Some are just used as audio and video peripherals and need to be connected to a display device or the terminal itself has an operating system.

[0007] Disadvantages of the existing technology:

[0008] Desktop microphones generally cannot directly run conference software and need to purchase an additional terminal or large screen with a system for use, which increases the complexity of system connection and the usage cost.

[0009] Conference terminal hosts also need to purchase peripherals to achieve the audio and video collection function, which increases the connection complexity and cost, and users cannot interact with the terminal intuitively and can only interact through a mouse, keyboard or remote control, making the wiring of meeting room devices messy and accessories easy to lose.

[0010] For conference large screens, since the large screen is generally placed at a relatively far distance from the meeting attendees, even if it is equipped with an array microphone, its effective pickup distance and pickup effect are not as good as those of an omnidirectional microphone placed on the desktop. And when users operate functions such as muting and volume adjustment, they need to get up and click on the screen, which is not as convenient as directly operating a desktop omnidirectional microphone on the desktop. The touch frame or touch screen adopted by the conference large screen to achieve interaction increases the usage cost of users.

[0011] The multimedia conference terminal has the same problems as the conference large screen in audio acquisition. Moreover, in terms of interaction, since there is no touch screen like the conference large screen, an external control tablet or a remote control needs to be used for control, which leads to problems such as poor interaction experience and easy loss of accessories.

[0012] Based on this, the present invention designs a split-type desktop conference terminal all-in-one machine to solve the above problems. Summary of the Invention

[0013] The purpose of the present invention is to propose a split-type desktop conference terminal all-in-one machine in order to solve the problems of poor interaction experience, poor sound pickup effect, complex equipment wiring, and easy loss and high cost of accessories in existing products.

[0014] In order to achieve the above purpose, the present invention adopts the following technical solutions:

[0015] A split-type desktop conference terminal all-in-one machine includes a base. Above the base, there is a conference host base. The base and the conference host base are connected by a rotating mechanism. On the side of the conference host base, 8 microphones are installed in an annular array for picking up the voices of speakers in the conference. A voice recognition system for controlling the 8 microphones to perform voice recognition is installed on the conference host base. The top of the conference host base is a touch tablet holder with magnetic charging, which is used to place the touch control tablet and charge it. The touch control tablet installs a self-developed application software based on the Android system, and realizes the control operation of the conference software through touch.

[0016] As a further description of the above technical solution:

[0017] The conference host base can rotate freely through the rotating mechanism, and the rotation angle does not exceed 360 degrees. At the bottom of the base, there are HDMI, USB, network ports, and power supply interfaces. The HDMI outputs the audio and video signals of the PC host to the conference display device, and the conference display device includes a TV, a projector, and a conference large screen.

[0018] As a further description of the above technical solution:

[0019] The conference host base contains a PC host with an X86 architecture. The PC host is equipped with a Windows system and is used to run the conference software and the self-developed conference control software. It has a high-fidelity speaker system and an 8-array microphone system.

[0020] As a further description of the above technical solution:

[0021] The touch control tablet is combined with the self-developed IOT device to control the relevant IOT devices in the meeting room. The touch control tablet supports handwriting and is used to project the handwriting content on the touch screen to the meeting display device in the meeting room. The touch control tablet comes with a built-in battery and can be used separately from the meeting host base.

[0022] As a further description of the above technical solution:

[0023] A front wide-angle camera is installed on the touch control tablet for video conferencing, and the front wide-angle camera can manually adjust the pitch angle.

[0024] As a further description of the above technical solution:

[0025] Upper magnetic charging contacts are provided at the bottom of the touch control tablet for charging the touch control tablet through the meeting host base, and it has a basic serial communication function. Corresponding lower magnetic charging contacts are provided on the meeting host base for charging the touch control tablet through the meeting host base, and it also has a basic serial communication function.

[0026] As a further description of the above technical solution:

[0027] The number of high-fidelity speakers is two, which are respectively distributed on the left and right sides to provide high-fidelity stereo sound effects.

[0028] As a further description of the above technical solution:

[0029] The voice recognition system includes an annular microphone array, an FPGA, a Pansy voice processing module, a voice recognition module, and a voice recognition output module;

[0030] The annular microphone array adopts an 8-microphone annular layout to perform voice algorithm processing on the recordings of 8 microphones, realizing voice interaction at a distance of 2-6 meters. The microphone has a built-in voice wake-up function and adopts a recording training method to improve the wake-up recognition rate, and can locate the position of the speaker, making the meeting host base turn to the speaker. The annular microphone array forms a pickup beam in the direction of the speaker, enhancing the speaker's voice and suppressing the surrounding background noise and reverberation;

[0031] The Pansy voice processing module adopts a specific hardware board with an annular layout.

[0032] As a further description of the above technical solution:

[0033] The voice recognition module is equipped with an offline voice recognition engine. The offline voice recognition engine uses the Lingyun offline vocabulary recognition technology. By loading the offline voice recognition engine and the offline voice package, it processes the voice for the local acoustic model and language model. The voice recognition module is connected to the offline recognition engine in a parallel manner. First, the keyword list is stored in the offline voice package of the recognition engine through the voice recognition module. The process of voice recognition is also the process of the voice recognition module completing its work. The text content recognized by the voice recognition module is matched with the keyword phrases in the list, and the keyword phrase with the highest score is found as the recognition result and output to the voice recognition output processing. The voice recognition output module plays the corresponding prompt sound.

[0034] As a further description of the above technical solution:

[0035] The voice recognition model is also equipped with a directional sound pickup problem model. The periodic change of sound pressure is caused during the sound transmission process. The microphone picks up sound by sensing the sound pressure through a sensor and converting it into an electrical signal. Assuming that the sound source signal at t 0 is h(n), then the frequency domain signal model received by the microphone at t i is:

[0036] Q(α,t i ,t 0 )=P(t i ,t 0 ,α)V(α,t 0 )L(α)S(α)+Z(α);

[0037] Among them,

[0038] is called the steady-state Green's function, which is the sound field excited by a steady-state point sound source located at t 0 , representing the delay and attenuation caused by the distance between the sound source and the microphone. V(α,t 0 ) represents the directivity of the microphone. Since the microphone is an omnidirectional microphone, its frequency response is usually flat in the range of 20 - 2000 Hz. Therefore, V(α,t 0 ) is usually a constant. L(α) is the frequency response of the amplifier and ADC. The frequency response of a high-performance amplifier should be flat in the passband. In addition, the ADC usually uses noise shaping technology to make the noise in the passband very small. Therefore, in most cases, it can be assumed that L(α)≡1. Z(α) is the noise, which usually includes two parts: correlated noise and uncorrelated noise.

[0039] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0040] 1. In the present invention, compared with the traditional split scheme of desktop microphones and conference terminals, this device integrates the functions of a conference host and a control screen on the basis of the capabilities of desktop microphones. Users can have a complete intelligent conference system experience without purchasing additional equipment.

[0041] 2. In the present invention, compared with traditional conference large screens, since the large screens are generally placed at a relatively far distance from the meeting participants, even if they are equipped with array microphones, their effective pickup distance and pickup effect are not as good as those of omnidirectional microphones placed on the desktop. Moreover, when users operate the mute or volume adjustment, they need to get up and click on the screen, which is not as convenient as directly operating on the desktop with a desktop omnidirectional microphone. The touch frames or touch screens adopted by conference large screens to achieve interaction increase the user's usage cost. The device of the present invention is placed on the desktop, which can solve the problems of too far pickup distance and inconvenient operation. And the device of the present invention has a touch screen, enabling the same operations as on the large screen while sitting in the seat.

[0042] 3. In the present invention, compared with traditional integrated multimedia conference terminal bars, there are the same problems in audio collection as those of conference large screens. And in terms of interaction, since there is no touch screen like that of conference large screens, an external control tablet or a remote control needs to be used for control, which brings problems such as poor interaction experience, complex wiring, and easy loss of accessories. The device of the present invention can also solve the problems of its usage experience.

[0043] 4. In the present invention, a pansy board with a circular layout is used as the core of voice front-end processing, combined with its corresponding offline speech recognition engine and language recognition module, and applied to the voice action control system of service robots. Through non-specific speech recognition tests at different distances, different angles, and echo cancellation in a noisy environment, the results show that in a noisy environment, the system has a high recognition rate for long-distance commands and can eliminate echoes, which is suitable for the application environment of service robots and also suitable for the application of far-field speech recognition systems in other noisy environments.

[0044] 5. In the present invention, the speech recognition model is equipped with a directional pickup problem model, which has a large amount of collected data, is stable and reliable in operation, has a very high transmission rate, is suitable for the directional pickup of multi-channel microphone arrays, and can provide assistance for the application of microphone arrays. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a schematic diagram of the overall structure of a split desktop conference terminal all-in-one machine proposed by the present invention;

[0046] Figure 2 It is a framework diagram of the speech recognition structure in a split desktop conference terminal all-in-one machine proposed by the present invention;

[0047] Figure 3The framework diagram of the voice recognition principle in a split desktop conference terminal all-in-one machine proposed by the present invention;

[0048] Figure 4 The signal reception model diagram of the microphone in a split desktop conference terminal all-in-one machine proposed by the present invention.

[0049] Legend description:

[0050] 1. Touch control tablet; 2. Touch tablet bracket; 3. Conference host base; 4. Array microphone; 5. Base; 6. Front wide-angle camera; 7. Upper magnetic charging contact; 8. Lower magnetic charging contact; 9. High-fidelity speaker. Specific implementation manners

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0052] Please refer to the attached Figure 1 - attached Figure 4 The present invention provides a technical solution: a split desktop conference terminal all-in-one machine, including a base 5. Above the base 5, a conference host base 3 is provided. The base 5 and the conference host base 3 are connected by a rotating mechanism. On the side of the conference host base 3, 8 microphones are installed in a circular array for picking up the voices of speakers in the conference. A voice recognition system for controlling the 8 microphones for voice recognition is installed on the conference host base 3. The top of the conference host base 3 is a touch tablet bracket 2 with magnetic charging for placing the touch control tablet 1 and charging it. The touch control tablet 1 installs a self-developed application software based on the Android system, and the control operation of the conference software is realized through touch.

[0053] As a further description of the above technical solution:

[0054] The conference host base 3 can rotate freely through the rotating mechanism, and the rotation angle does not exceed 360 degrees. At the bottom of the base 5, HDMI, USB, network ports, and power supply interfaces are provided. The HDMI outputs the audio and video signals of the PC host to the conference display device, and the conference display device includes a TV, a projector, and a conference large screen.

[0055] As a further description of the above technical solution:

[0056] The conference host base 3 includes an X86-based PC host with a Windows system, which is used to run conference software and self-developed conference control software. It has a high-fidelity speaker 9 system and an 8-array microphone 4 system.

[0057] As a further description of the above technical solution:

[0058] The touch control panel 1 is combined with self-developed IOT devices to control the relevant IOT devices in the conference room. The touch control panel 1 supports handwriting and is used to project the handwritten content on the touch screen to the conference display device in the conference room. The touch control panel 1 has a built-in battery and can be used separately from the conference host base 3.

[0059] As a further description of the above technical solution:

[0060] The touch control panel 1 is equipped with a front wide-angle camera 6 for video conferencing, and the front wide-angle camera 6 can manually adjust the pitch angle.

[0061] As a further description of the above technical solution:

[0062] The bottom of the touch control panel 1 is provided with upper magnetic charging contacts 7 for charging the control panel through the conference host base 3 and with basic serial communication function. Corresponding to the upper magnetic charging contacts 7 on the conference host base 3, there are lower magnetic charging contacts 8 for charging the touch control panel 1 through the conference host base 3 and with basic serial communication function.

[0063] As a further description of the above technical solution:

[0064] The number of high-fidelity speakers 9 is two, which are distributed on the left and right sides respectively to provide high-fidelity stereo sound effects.

[0065] As a further description of the above technical solution:

[0066] The voice recognition system includes a ring microphone array, an FPGA, a Pansy voice processing module, a voice recognition module, and a voice recognition output module;

[0067] The ring microphone array adopts an 8-microphone circular layout to perform voice algorithm processing on the recordings of 8 microphones, realizing voice interaction at a distance of 2-6 meters. The microphone has a built-in voice wake-up function and uses a recording training method to improve the wake-up recognition rate. It can also locate the speaker's position, causing the conference host base 3 to turn towards the speaker. The ring microphone array forms a pickup beam in the direction of the speaker, enhancing the speaker's voice and suppressing surrounding background noise and reverberation;

[0068] The Pansy voice processing module uses a specific hardware board with a circular layout.

[0069] As a further description of the above technical solution:

[0070] The voice recognition module is equipped with an offline voice recognition engine. The offline voice recognition engine uses Lingyun offline vocabulary recognition technology. By loading the offline voice recognition engine and the offline voice package, it processes the voice for the local acoustic model and language model. The voice recognition module is connected to the offline recognition engine in a parallel manner. First, the keyword list is stored in the offline voice package of the recognition engine through the voice recognition module. The process of voice recognition is also the process of the voice recognition module completing its work. It matches the text content recognized by the voice recognition module with the keyword phrases in the list, finds the keyword phrase with the highest score as the recognition result, and outputs it to the voice recognition output processing. The voice recognition output module plays the corresponding prompt sound.

[0071] As a further description of the above technical solution:

[0072] The voice recognition model is also equipped with a directional sound pickup problem model. The sound pressure changes periodically during the sound transmission process. The microphone picks up the sound by sensing the sound pressure through the sensor and converting it into an electrical signal. Assuming that the sound source signal at t 0 is h(n), then the frequency-domain signal model received by the microphone at t i is:

[0073] Q(α,t i ,t 0 )=P(t i ,t 0 ,α)V(α,t 0 )L(α)S(α)+Z(α);

[0074] Among them,

[0075] is called the steady-state Green's function, which is the sound field excited by a steady-state point source at t 0 , representing the delay and attenuation caused by the distance between the sound source and the microphone. V(α,t 0 ) represents the directivity of the microphone. Since the microphone is an omnidirectional microphone, its frequency response is usually flat in the range of 20 - 2000 Hz. Therefore, V(α,t 0 ) is usually a constant. L(α) is the frequency response of the amplifier and ADC. The frequency response of a high-performance amplifier should be flat in the passband. In addition, the ADC usually uses noise shaping technology to make the noise in the passband very small. Therefore, in most cases, it can be assumed that L(α)≡1. Z(α) is the noise, usually including two parts: correlated noise and uncorrelated noise.

[0076] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention should cover within the protection scope of the present invention any equivalent replacement or change made according to the technical solution and inventive concept of the present invention.

Claims

1. A split-type desktop conference terminal, comprising a base (5), characterized in that: A conference host base (3) is arranged above the base (5), and the base (5) is connected to the conference host base (3) via a rotating mechanism. Eight microphones in a circular array are installed on the side of the conference host base (3) for picking up the voice of the speaker in the meeting. A voice recognition system for controlling the eight microphones for voice recognition is installed on the conference host base (3). The top of the conference host base (3) is a touch tablet bracket (2) with magnetic charging, which is used to place and charge the touch conference control tablet (1). The touch conference control tablet (1) is installed with self-developed application software based on the Android system, and the control operation of the conference software is realized by touch.

2. The split-type desktop conference terminal according to claim 1 is characterized in that: The conference host base (3) is freely rotatable via a rotating mechanism, and the rotation angle does not exceed 360 degrees. The bottom of the base (5) is provided with an HDMI, USB, network port and power supply interface. The HDMI outputs the audio and video signals of the PC host to the conference display device. The conference display device includes a television, a projector and a large conference screen.

3. The split-type desktop conference terminal according to claim 1 is characterized in that: The conference host base (3) includes a PC host with an X86 architecture, the PC host is equipped with a Windows system and is used to run conference software and self-developed conference control software, and has a high-fidelity speaker (9) system and an 8-array microphone (4) system.

4. The split-type desktop conference terminal according to claim 1 is characterized in that: The touch conference control tablet (1) is combined with a self-developed IOT device to realize the control operation of the conference room-related IOT device. The touch conference control tablet (1) supports handwriting and is used to project the handwritten content on the touch screen to the conference display device in the conference room. The touch conference control tablet (1) has its own battery and is used separately from the conference host base (3).

5. The split-type desktop conference terminal according to claim 1 is characterized in that: The touch conference control tablet (1) is equipped with a front wide-angle camera (6) for video conferencing, and the front wide-angle camera (6) can manually adjust the pitch angle.

6. The split-type desktop conference terminal according to claim 1, characterized in that: The bottom of the touch conference control tablet (1) is provided with an upper magnetic charging contact (7) for charging the conference control tablet via the conference host base (3) and having a basic serial communication function; the conference host base (3) is provided with a lower magnetic charging contact (8) corresponding to the upper magnetic charging contact (7) for charging the touch conference control tablet (1) via the conference host base (3) and having a basic serial communication function.

7. The split-type desktop conference terminal according to claim 3 is characterized in that: The number of the high-fidelity loudspeakers (9) is two, which are respectively distributed on the left and right sides, and are used to provide high-fidelity stereo sound effects.

8. The split-type desktop conference terminal according to claim 1, characterized in that: The speech recognition system includes a ring microphone array, an FPGA, a Pansy speech processing module, a speech recognition module, and a speech recognition output module; The circular microphone array adopts an 8-microphone circular layout to process the recordings of the 8-channel microphones with a voice algorithm to achieve 2-6 meters long-distance voice interaction. The microphone has a voice wake-up function and adopts a recording training method to improve the wake-up recognition rate. It can also locate the speaker's position, so that the conference host base (3) is turned towards the speaker. The circular microphone array forms a sound pickup beam in the direction of the speaker, enhances the speaker's voice, and suppresses the surrounding background sound and reverberation. The Pansy voice processing module uses a specific hardware board in a ring layout.

9. The split-type desktop conference terminal according to claim 8, characterized in that: The speech recognition module is equipped with an offline speech recognition engine, which uses Lingyun offline vocabulary recognition technology. By loading the offline speech recognition engine and the offline speech package, the speech is processed with a localized acoustic model and a language model. The speech recognition module is connected to the offline recognition engine in parallel. The keyword list is first stored in the offline voice package of the recognition engine through the speech recognition module. The speech recognition process is also the process of the speech recognition module completing its work. The text content recognized by the speech recognition module is matched with the key words in the list, and the keyword with the highest score is found as the recognition result and input into the speech recognition output processing. The speech recognition output module plays the corresponding prompt sound.

10. The split-type desktop conference terminal according to claim 9, characterized in that: The speech recognition model is also equipped with a directional sound pickup problem model. The sound transmission process causes periodic changes in sound pressure. The microphone senses the sound pressure through the sensor and converts it into an electrical signal for sound pickup. Assuming that the sound source signal at t0 is h(n), then the sound source signal at t i The frequency domain signal model received by the microphone at is: Q(α,t i ,t0)=P(t i ,t0,α)V(α,t0)L(α)S(α)+Z(α); in, It is called the steady-state Green's function, which is the sound field excited by the steady-state point sound source at t0. It represents the delay and attenuation caused by the distance between the sound source and the microphone. V(α, t0) represents the directivity of the microphone. Since the microphone is an omnidirectional microphone, its frequency response is usually flat in the range of 20-2000Hz, so V(α, t0) is usually a constant. L(α) is the frequency response of the amplifier and ADC. The frequency response of a high-performance amplifier in the passband should be flat. In addition, the ADC usually uses noise shaping technology to make the noise in the passband very small. Therefore, in most cases, it can be assumed that L(α)≡1. Z(α) is the noise, which usually includes two parts: correlated noise and uncorrelated noise.