Restaurant system and method based on Internet real-time audio and video interconnection

By optimizing equipment configuration and transmission mechanisms, and combining professional equipment layout and diversified interconnection modes, the problems of unreasonable equipment configuration, poor audio-visual synchronization and insufficient immersion in communication between restaurants in different locations have been solved, realizing real-time, high-definition, and low-latency audio-visual interconnection between local and remote restaurant private rooms.

CN120935320APending Publication Date: 2025-11-11周玉松
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511090245.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies for communication in restaurants in different locations suffer from problems such as unreasonable equipment configuration, poor audio-visual synchronization, single interconnection mode, and insufficient immersion. This results in low resolution and unstable frame rate of video call equipment, limited sound pickup range of audio equipment, and high transmission latency, making it impossible to meet the needs of simultaneous interconnection in multiple locations.

Method used

Employing high-resolution and high-refresh-rate visual display devices, a circular array microphone and multi-channel audio, a high-definition camera, and an internet communication module, combined with professional equipment configuration, precise positioning, and an efficient transmission mechanism, it achieves a closed-loop audio-visual synchronization and supports diverse interconnection modes.

Benefits of technology

It enables real-time, high-definition, low-latency audio-visual interconnection between local and remote restaurant private rooms, improving the rationality of equipment configuration, audio-visual synchronization, and immersive experience, and adapting to the needs of different dining sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120935320A_ABST
    Figure CN120935320A_ABST
Patent Text Reader

Abstract

The invention discloses a restaurant system and method based on Internet real-time audio and video interconnection, and the system comprises interconnection terminals of local and remote restaurants, and each terminal comprises a visual display device, an audio playing device, a visual collection device, an audio collection device, an Internet communication module and a dining table. The visual display equipment adopts a high-resolution display device and is deployed on a decorated wall surface, a curtain projection, or in front of a dining table, or embedded into the dining table and the like; the audio equipment comprises a high-fidelity microphone and a multi-channel sound box and is used for accurately collecting and playing sound; the visual acquisition device is a high-definition camera and is installed by being matched with a suspended ceiling or the side edge of a screen. The internet communication module guarantees low-delay sound and picture transmission based on the high-speed transmission characteristic of the internet, and two-way synchronization is formed. The system supports three interconnection modes of mirroring, complementation and single point, is respectively adapted to single-remote, multi-remote and master-slave interconnection scenes, can effectively optimize the problems of asynchronous sound and picture, poor experience and the like in remote dinner party, and improves the immersion and flexibility of remote interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of real-time audio-visual interaction, and in particular to a restaurant system and method based on real-time audio-visual interconnection via the Internet. Background Technology

[0002] In the internet age, the demand for gatherings and communication among relatives, friends, and business partners in different locations is growing, but geographical distance makes traditional face-to-face dining difficult to achieve. While existing technologies such as video calls and online meetings can be used for remote communication, they have significant limitations in the catering industry context:

[0003] 1. Insufficient equipment configuration: Existing video call equipment (such as ordinary cameras and home TVs) has low resolution and unstable frame rate, which cannot clearly present the details of the dining scene; audio equipment (such as ordinary microphones and speakers) has limited sound pickup range and weak noise resistance, which is prone to sound distortion or echo, affecting the communication experience.

[0004] 2. Unreasonable location layout: The existing equipment lacks professional deployment for restaurant scenarios. The distance and angle between the display equipment and the table are arbitrary, resulting in uncomfortable viewing angles for viewers. The position of the acquisition equipment is deviated, and it cannot fully cover the table and the area where diners are, resulting in fragmented or missing images.

[0005] 3. Poor real-time synchronization: Current Internet transmission mostly uses general protocols that are not optimized for audio-visual synchronization, resulting in high transmission latency (often exceeding 200ms), obvious audio-visual asynchrony, and difficulty in creating an immersive experience.

[0006] 4. Limited interconnection modes: Existing technologies mostly support simple one-to-one video connections, which cannot meet the needs of simultaneous interconnection in multiple locations, and lack adaptability to modes for different gathering sizes, resulting in insufficient flexibility.

[0007] To address the aforementioned issues, this invention proposes a restaurant system based on real-time audio-visual interconnection via the Internet. Through professional equipment configuration, precise location layout, efficient transmission mechanism, and diverse interconnection modes, it achieves high-quality, low-latency real-time audio-visual interconnection between private rooms in restaurants located in different places, thus overcoming the shortcomings of existing technologies. Summary of the Invention

[0008] This application aims to at least partially address one of the technical problems in the related art.

[0009] Therefore, one objective of this application is to provide a restaurant system and method based on real-time audio-visual interconnection via the Internet, which solves the technical problems existing in the current methods of communication between restaurants in different locations, such as unreasonable equipment configuration, poor audio-visual synchronization, single interconnection mode, and insufficient immersion, and provides a system that can realize real-time, high-definition, and low-latency audio-visual interconnection between local and remote restaurant private rooms.

[0010] To achieve the above objectives, the first aspect of this application proposes a restaurant system based on real-time audio-visual interconnection via the Internet, comprising interconnected terminals deployed in a local restaurant and at least one remote restaurant. Each interconnected terminal includes a visual display device, an audio device, a visual acquisition device, an Internet communication module, and a dining table. The visual display device employs at least one of an LED screen, a television screen, and a projection device, with a resolution of at least 4K UHD (3840×2160 pixels), a refresh rate of at least 60Hz, and is installed on the wall directly in front of the dining table or embedded within the table, at a distance of 1.2-2.0 meters from the table. The audio device includes a ring array microphone and multi-channel surround sound. The microphone is installed at the center of the dining table, comprising 6-8 directional microphone units, a sampling rate of at least 48kHz, and a signal-to-noise ratio of at least 70dB. The speakers are symmetrically arranged on both sides of the visual display device, with an output power of 80-120W and a frequency response range of 20Hz-20kHz. The visual acquisition device is a high-definition camera with a resolution of ≥1080p, a frame rate of ≥30fps, and a wide-angle lens with a field of view of ≥120°. It is positioned either directly above the ceiling or on the side bracket of the visual display device, at a height of 2.0-2.5 meters from the table surface. The Internet communication module is based on the TCP / IP protocol stack, with a built-in WebRTC real-time transmission protocol, supporting H.265 video encoding and Opus audio encoding, a transmission bandwidth of ≥10Mbps / channel, and an end-to-end latency of ≤100ms. The output of the visual acquisition device in the local restaurant is connected to the input of the Internet communication module, and the output of the Internet communication module in the remote restaurant is connected to the input of the local visual display device and the speakers, forming a two-way audio-visual synchronous closed loop.

[0011] According to an embodiment of this application, a restaurant system and method based on real-time audio-visual interconnection via the Internet is proposed to solve the technical problems existing in the current methods of communication between restaurants in different locations, such as unreasonable equipment configuration, poor audio-visual synchronization, single interconnection mode, and insufficient immersion. The system provides a system that can realize real-time, high-definition, and low-latency audio-visual interconnection between local and remote restaurant private rooms.

[0012] In addition, the restaurant system and method based on real-time audio-visual interconnection via the Internet proposed in this application may also have the following additional technical features:

[0013] In one embodiment of this application, the visual display device supports HDR high dynamic range display, has a screen size of 55-85 inches, an adjustable installation angle (tilt angle ±15°), and an anti-glare coating on the screen surface.

[0014] In one embodiment of this application, the diameter of the circular array of microphones is 0.5-0.8 meters, the included angle between adjacent microphone units is 45°-60°, and beamforming technology is used to directionally collect human voices; the sound system is a 5.1 channel system, with the center channel aligned with the center of the dining table.

[0015] In one embodiment of this application, the visual acquisition device has a built-in PTZ gimbal mechanism with a horizontal rotation range of ±90° and a vertical rotation range of ±30°, and automatically tracks and locates diners through a face recognition algorithm.

[0016] In one embodiment of this application, the Internet communication module includes a main server and an edge computing node. The main server is deployed in the cloud and is used to calibrate remote audio and video timestamps. The edge computing node is deployed on a local router and is used to compress data packets in real time with a compression rate of 50%-60%.

[0017] In one embodiment of this application, the positional relationship of each device is as follows: the center point of the visual display device and the center point of the dining table are on the same vertical plane; the center point of the microphone array coincides with the geometric center of the dining table, and the height is 0.2-0.3 meters above the tabletop; the optical axis of the visual acquisition device is perpendicular to the surface of the dining table, and the downward angle is 10°-20°.

[0018] In one embodiment of this application, the main router of the Internet communication module is deployed in the restaurant equipment room and is directly connected to the visual display device, audio and visual acquisition device via gigabit network cable. Wireless transmission is only used for communication between the array units of the microphones.

[0019] In one embodiment of this application, the system uses a mirroring mode where a local restaurant is interconnected with a single remote restaurant. The local visual display device is divided into two display areas: left and right. The left area displays the complete image captured by the remote restaurant's visual acquisition device in real time, while the right area displays local auxiliary information. The audio stream captured by the remote microphone is output through the left channel of the local speaker, with a 3-5ms buffer added during output to align with the image. The wide-angle lens of the visual acquisition device covers the entire remote dining table area, and the image is transmitted without cropping.

[0020] In one embodiment of this application, the system's completion mode allows the local restaurant to simultaneously connect to at least two remote restaurants. The visual display device is divided into a 2×2 grid area; each grid independently displays the image of one remote restaurant with a 16:9 aspect ratio and adaptive resolution scaling; the audio device employs a dynamic mixing algorithm, with each remote sound source independently mapped to different speaker azimuth angles; the internet communication module enables a multi-point relay server with dynamic bandwidth allocation (single channel ≥ 5Mbps).

[0021] In one embodiment of this application, the system operates in a single-point mode, connecting only a single restaurant in a different location, with 85%-90% of the visual display area showing the scene from that location; the visual acquisition device is fixed in wide-angle mode and has no PTZ tracking function; the internet communication module's transmission bandwidth is reduced to 5-8Mbps, with a latency of ≤120ms; the audio is simplified to stereo output, and the microphone array units are reduced to 4.

[0022] The advantages of this application compared to existing technologies are:

[0023] (1) By optimizing the resolution, refresh rate and installation position of the visual display device, combined with anti-glare coating and angle adjustment design, the clear display of the screen in different locations and the comfortable viewing experience are ensured, and the problems of blurry screen and uncomfortable viewing of the existing device are solved.

[0024] (2) The audio equipment uses a ring array microphone and multi-channel speakers, combined with beamforming technology and precise positioning layout, which improves the directionality of sound acquisition and the stereo effect of playback, effectively reduces noise and echo interference, and improves the quality of remote audio interaction.

[0025] (3) The high-definition parameters, wide-angle lens and tracking function of the visual acquisition equipment, combined with reasonable installation height and angle, can completely and accurately collect local dining scenes, providing a high-quality image source for real-time presentation in different locations.

[0026] (4) The Internet communication module adopts a dedicated protocol and encoding method, combined with cloud calibration and edge computing, to achieve low-latency transmission and synchronization of audio and video, with end-to-end latency controlled within 100ms, thus solving the core problem of poor synchronization in the existing technology.

[0027] (5) The three interconnection modes are designed for different numbers of restaurants in different locations, flexibly adapting to one-to-one, many-to-many dining needs, expanding the scope of application of the system, and are more practical than the existing single mode.

[0028] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0029] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0030] Figure 1 This is a schematic diagram of system connection control for a restaurant system and method based on real-time audio-visual interconnection via the Internet, according to an embodiment of this application.

[0031] Figure 2This is a schematic diagram illustrating the overall electrical connections and control relationships of a restaurant system and method based on real-time audio-visual interconnection via the Internet, according to an embodiment of this application.

[0032] Figure 3 This is a diagram showing the device location relationships and hardware connection topology of a restaurant system and method based on real-time audio-visual interconnection via the Internet, according to an embodiment of this application.

[0033] Figure 4 A flowchart illustrating the mirror application mode control of a restaurant system and method based on real-time audio-visual interconnection via the Internet, according to an embodiment of this application.

[0034] Figure 5 This is a supplementary application mode data flow diagram of a restaurant system and method based on real-time audio-visual interconnection over the Internet according to an embodiment of this application.

[0035] As shown in the figure: 1. Visual display device; 2. Audio device; 3. Visual acquisition device; 4. Internet communication module; 5. Dining table; 21. Microphone; 22. Speaker; 41. Main server; 42. Edge computing node. Detailed Implementation

[0036] Embodiments of this application are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application. Rather, embodiments of this application include all variations, modifications, and equivalents falling within the spirit and scope of the appended claims.

[0037] The following description, in conjunction with the accompanying drawings, illustrates a restaurant system and method based on real-time audio-visual interconnection via the Internet, representing an embodiment of this application.

[0038] like Figures 1-5 As shown in the embodiment of this application, a restaurant system and method based on real-time audio-visual interconnection via the Internet is described below:

[0039] Example 1: This system achieves real-time audio-visual interconnection between private dining rooms in different locations through the collaborative work of visual display device 1, audio device 2, visual acquisition device 3, and internet communication module 4. The specific configuration and deployment of each component are as follows:

[0040] 1. Visual display equipment

[0041] Equipment selection: 65-inch Samsung QN65Q80C 4K LED screen (resolution 3840×2160 pixels, refresh rate 60Hz), supports HDR10 display, and the screen surface is covered with AG anti-glare coating (haze value 20%).

[0042] Installation requirements:

[0043] The screen is mounted on the wall directly in front of the dining table 5, with the center point of the screen and the center point of the dining table 5 on the same vertical plane (deviation ≤ 5cm). The bottom edge of the screen is 1.3 meters from the ground and 1.5 meters horizontally from the edge of the dining table 5 (within the range of 1.2-2.0 meters).

[0044] Small private rooms can be embedded inside the dining table 5, with an embedding depth of 1 / 2 the thickness of the screen, ensuring that the screen surface is flush with the tabletop and avoiding obstruction of the view during meals.

[0045] 2. Audio equipment

[0046] Circular array microphone:

[0047] The array consists of 21 Audio-Technica AT9904 directional microphones with 8 units arranged in a circular array (0.6 meters in diameter), with an angle of 45° between adjacent units, a sampling rate of 48kHz, and a signal-to-noise ratio of 70dB.

[0048] Installed 0.25 meters directly above the geometric center of the dining table 5 (0.2-0.3 meters above the tabletop), and fixed with a bracket. The cable is threaded through the pre-drilled hole in the center of the dining table 5 and then under the table to avoid interfering with dining.

[0049] Multi-channel surround sound:

[0050] The JBL Cinema 510 5.1 channel system (output power 100W, frequency response range 20Hz-20kHz) was selected, with the center channel facing the center of the dining table 5, and the left and right channels symmetrically arranged on both sides of the visual display device 1 (2.0 meters apart).

[0051] A 2cm thick sound-absorbing cotton is placed between the speaker 22 and the microphone 21 as an acoustic isolation barrier to reduce echo interference.

[0052] 3. Visual acquisition equipment

[0053] Equipment configuration: Uses Hikvision DS-2CD3T27WD-L 2-megapixel camera (1080p resolution, 30fps frame rate), equipped with a 120° wide-angle lens;

[0054] Installation location:

[0055] Ceiling installation: 2.2 meters from the surface of dining table 5 (within the range of 2.0-2.5 meters), with the optical axis perpendicular to the tabletop and a downward angle of 15° (within the range of 10°-20°), ensuring that the lens covers dining table 5 and the surrounding 1.5-meter area;

[0056] Alternative solution: Fix the lens to the side bracket of the visual display device 1 (1.7 meters from the ground), with the lens pointing horizontally towards the center of the dining table 5 to avoid incomplete images due to shooting angle deviation.

[0057] It should be noted that the face recognition and tracking implementation is as follows:

[0058] The visual acquisition device 3 has a built-in CNN (Convolutional Neural Network)-based facial recognition model, specifically a lightweight MobileNet-SSD architecture. The input resolution is 640×480 pixels (processing time per frame ≤30ms), and the output is the coordinates (x,y) of the center point of the diner's face and the bounding box size (width×height).

[0059] Tracking process: When facial movement is detected, the PTZ gimbal mechanism adjusts in real time according to the coordinate deviation (horizontal / vertical rotation speed is proportional to the deviation value, maximum 30° / s) to ensure that the face is always in the center area of ​​the image (deviation ≤ 5cm);

[0060] Multi-person tracking: Supports simultaneous recognition of 3-5 faces, uses non-maximum suppression (NMS) algorithm to remove duplicate detection boxes, and prioritizes tracking the most active target (determined by blinking, head turning, etc.).

[0061] 4. Internet communication module

[0062] Hardware components:

[0063] Edge computing node 42: It adopts Huawei Honor Router Pro2, with built-in H.265 video encoding module (compression rate of 55%), and is deployed in the restaurant equipment room;

[0064] Master server 41: Tencent Cloud CVM instance (select the node closest to the user), running NTP time synchronization service to ensure that audio and video timestamps are consistent across different locations;

[0065] Transmission Protocol: Based on the TCP / IP protocol stack, with built-in WebRTC real-time transmission protocol, supporting H.265 video encoding and Opus audio encoding, single-channel transmission bandwidth of 10Mbps, end-to-end latency ≤100ms, realizing the core function of "real-time synchronization of video and audio across different locations via the Internet".

[0066] II. System Workflow

[0067] 1. Startup and initialization (time ≤ 3s)

[0068] Device self-test: Visual acquisition device 3 starts automatic white balance calibration, microphone 21 collects ambient noise to establish a baseline value, and network module tests uplink / downlink bandwidth (must be ≥30Mbps / 50Mbps respectively);

[0069] Connection establishment: The local terminal sends its device identifier to the main server 41, obtains the IP address of the remote terminal, exchanges media description information through the WebRTC protocol, and establishes a P2P connection.

[0070] 2. Real-time audio and video transmission closed loop (delay ≤ 100ms)

[0071] Local video and audio capture:

[0072] 1. Visual acquisition device 3 captures one frame every 33ms (30fps), and adds an NTP timestamp after H.265 encoding (bitrate 8Mbps);

[0073] 2. The microphone array 21 uses beamforming technology to directionally collect human voices within 1.5 meters, and adds timestamps synchronously after Opus encoding (256kbps);

[0074] 3. The audio and video streams are multiplexed into a single data packet at the edge computing node 42 and transmitted to the main server 41 via a gigabit network cable.

[0075] Remote audio and video reception and output:

[0076] 4. The main server 41 calibrates the timestamp of remote data packets (error ≤ 5ms) and forwards them to the local Internet communication module 4;

[0077] 5. The local edge node demultiplexes the data packets, sends the video stream to the visual display device 1 for decoding and display, and adds a 3ms buffer (to compensate for decoding delay) to the audio stream before outputting it through the speaker 22;

[0078] 6. Local audio and video are transmitted to remote terminals using the same process, forming a two-way synchronous closed loop of "local acquisition → remote output → local reception".

[0079] III. Specific Implementation of the Three Application Modes

[0080] 1. Mirror application mode

[0081] Applicable scenario: Two-way interconnection between private rooms in city A and private rooms in city B.

[0082] Equipment matching: The terminal equipment models and installation locations in the two locations are completely symmetrical (e.g., the camera in city A and the camera in city B have the same height and angle);

[0083] The screen displays: 70% of the left side of the local visual display device 1 shows the complete view from the other location (120° wide-angle coverage by the camera), and 30% of the right side shows the local time and menu information;

[0084] Audio processing: Audio from other locations is output through the local left channel, and local audio is output through the left channel of other locations. A 3ms buffer ensures audio-visual synchronization, creating an immersive feeling of "sitting opposite each other".

[0085] 2. Complete the application mode

[0086] Applicable scenario: A private room in city A is connected to two other private rooms in cities B and C.

[0087] Screen layout: The A city private room is divided into a 2×2 grid (each grid is 1920×1080 pixels) as the local visual display device. The B city private room screen is displayed in the upper left corner, the C city private room screen is displayed in the upper right corner, and the local preview is displayed in the lower left corner.

[0088] Sound processing: A dynamic mixing algorithm is used. The sound from the B city room is mapped to the left front channel (azimuth angle -30°), and the sound from the C city room is mapped to the right front channel (azimuth angle +30°). The source of the speech is distinguished by sound field localization.

[0089] Network optimization: Enable multi-point relay servers, allocate a minimum bandwidth of 5Mbps to each of the two connections, and keep the total bandwidth usage ≤15Mbps to avoid congestion.

[0090] 3. Single-point application mode

[0091] Applicable scenario: Interconnection between the main private room in city A and the home terminal in city D.

[0092] Simplified equipment: The home terminal uses a 55-inch Xiaomi TV (replacing the LED screen), 4-unit microphones 21 (array diameter 0.5 meters), and 2.0 channel speakers 22 (output power 20W);

[0093] The screen displays: 88% of the area of ​​the main private room visual display device 1 in City A displays the home screen, and 12% of the area displays the home network status (bandwidth, latency);

[0094] Parameter adjustments: Transmission bandwidth reduced to 6Mbps, latency tolerance limit 120ms, speaker 22 simplified to stereo output, adapted to home network environment.

[0095] One point that needs further explanation is the bandwidth adaptive adjustment mechanism:

[0096] The home terminal's internet communication module 4 checks the real-time uplink bandwidth every 500ms. When the measured value is less than 10Mbps for three consecutive times, a degradation process is automatically triggered.

[0097] 1. Edge computing node 42 reduces the video encoding bitrate from 10Mbps to 6Mbps (H.265 CRF value adjusted from 23 to 28), maintaining the resolution of 1080p but reducing the I-frame interval (from 30 frames to 60 frames);

[0098] 2. The audio encoding bitrate was reduced from 256kbps to 128kbps, the frame length was increased from 20ms to 40ms, and the number of data packets was reduced;

[0099] 3. Visual acquisition device 3 disables PTZ tracking function, fixes wide-angle mode to reduce computing power consumption, and turns off autofocus (uses fixed focal length 50cm-3m);

[0100] Once the bandwidth recovers to ≥10Mbps and remains stable for 3 seconds, the system automatically reverts to standard parameters, ensuring that home terminals can maintain a basic interactive experience even when the network fluctuates.

[0101] IV. Commissioning and Maintenance Specifications

[0102] 1. Installation and debugging:

[0103] The center deviation between the visual display device 1 and the dining table 5 needs to be calibrated with a laser level (≤5mm);

[0104] Audio-visual synchronization test: Play a standard test video (including stopwatch and beat sound), measure the delay with an oscilloscope, and ensure it is ≤20ms.

[0105] 2. Routine maintenance:

[0106] Daily checks include the cleanliness of the camera lens and the sensitivity of the microphone 21.

[0107] Test network latency weekly using specialized tools. If latency exceeds 100ms, contact your ISP to optimize the line.

[0108] 3. Troubleshooting:

[0109] When the network is interrupted, the edge nodes automatically cache data for 5 seconds and quickly synchronize after reconnection;

[0110] When microphone 21 malfunctions, the system automatically switches to the backup pickup channel (microphone 21 is reserved in the corner of the dining table 5) to ensure that the call is not interrupted.

[0111] To clearly illustrate the above embodiments, in the completion application mode, audio device 2 uses a dynamic mixing algorithm based on azimuth mapping to achieve sound source separation, and the specific formula is as follows:

[0112] Let the signal from the sound source in a different location be S. i (t)(i=1,2,…,n,n is the number of restaurants in different locations), the corresponding mapped azimuth angle of the sound 22 is θ. i (Unit: degrees, range: -90° to 90°), then the local left channel output signal L(t) and right channel output signal R(t) are respectively:

[0113]

[0114] The cosine function is used to allocate the gain coefficients (range 0-1) of the left and right channels according to the azimuth angle. For example, when θ i =0° (directly above), left / right channel gain is 1, signal is output simultaneously; when θ i =90° (directly to the right), the left channel gain is 0, the right channel gain is 1, and the signal is output only from the right channel, achieving accurate differentiation of the sound source location.

[0115] When tracking multiple faces, visual acquisition device 3 uses the non-maximum suppression (NMS) algorithm to filter duplicate detection boxes. The specific parameters are as follows:

[0116] IoU (Intersection over Union) threshold: set to 0.35. The calculation formula is:

[0117]

[0118] When the IoU of two bounding boxes is ≥ 0.35, they are determined to be the same target. The bounding box with higher confidence is retained, and the overlapping bounding boxes are deleted.

[0119] Confidence threshold: Set to 0.6. Detection boxes with a value lower than this will be filtered out directly to reduce false detections.

[0120] This parameter setting is adapted to the characteristics of restaurant scenarios where diners are close together (the distance between faces may be <50cm), ensuring tracking accuracy while avoiding missing adjacent faces.

[0121] It should be noted that the control method of this application can be automatically controlled by a controller. The control method of the controller can be implemented by simple programming by those skilled in the art, which is common knowledge in the field. Furthermore, this application is mainly used to protect mechanical structures, so the control method and circuit connection will not be explained in detail here.

[0122] Specifically, in actual implementation, taking the real-time group dining achieved through mirroring between private rooms in City A (local) and private rooms in City B (remote) as an example, the complete workflow is as follows:

[0123] 1. Equipment Startup and Initialization: After the user enters the private room in City A, the staff turns on the main power of the system, the visual display device 1 lights up for self-test, and the screen displays the initialization interface; the visual acquisition device 3 automatically turns on, the lens completes focusing and starts the MobileNet-SSD face recognition model to capture diners in the area of ​​the dining table 5 in real time; the ring array microphone 21 starts to collect environmental noise and establish a baseline value; the main server 41 and edge computing node 42 of the Internet communication module 4 start up, test the network bandwidth of A and B (both ≥30Mbps), exchange signaling through the WebRTC protocol, and establish a P2P connection (takes about 2 seconds).

[0124] 2. Remote Connection and Synchronization: The visual acquisition device 3 in the private room of City A captures one frame (1080p resolution) every 33ms, encodes it with H.265, adds an NTP timestamp, and transmits it to the Internet communication module 4. Simultaneously, the circular array microphone 21 collects the conversations of diners in City A, encodes them with Opus, adds a timestamp, and multiplexes them with the video stream into a data packet at the edge computing node 42. This data packet is transmitted to the main server 41 via a gigabit network cable, and after the timestamp is calibrated, it is forwarded to the Internet communication module 4 in the private room of City B.

[0125] 3. Remote Video and Audio Output: After receiving data packets, the internet communication module 4 in the private room of City B demultiplexes them into video and audio streams by the edge computing node 42. The video stream is sent to the visual display device 1 in City B, where it decodes and displays the real-time video of City A in the left 70% area (because the visual acquisition device 3 uses a 120° wide-angle lens, it fully covers the dining table 5 and its surroundings); the audio stream is buffered by 3ms by the pre-processor of the multi-channel surround sound speaker 22 and output through the left channel to ensure synchronization with the video.

[0126] 4. Local Reception and Interaction: The visual acquisition device 3 and the circular array microphone 21 in the private room of City B acquire audio and video through the same process and transmit them to the private room of City A via the Internet communication module 4. The visual display device 1 in City A displays the image of City B in the left 70% area, and the multi-channel surround sound 22 outputs the sound of City B in the left channel, forming a "face-to-face" interaction. During this time, if the diners in City A move, the PTZ pan-tilt unit of the visual acquisition device 3 adjusts its angle in real time according to the facial recognition coordinates to ensure that they are always in the center of the image (deviation ≤ 5cm).

[0127] 5. End and Shutdown: After the dinner, the staff shuts down the system, the Internet communication module 4 disconnects the P2P connection, and the visual display device 1, visual acquisition device 3, and audio device 2 enter standby mode in sequence, completing a complete remote real-time audio-visual interconnection experience.

[0128] In summary, the restaurant system and method based on real-time audio-visual interconnection via the Internet, as described in this application, solves the technical problems existing in current cross-regional restaurant communication methods, such as unreasonable equipment configuration, poor audio-visual synchronization, single interconnection mode, and insufficient immersion. It provides a system that can achieve real-time, high-definition, and low-latency audio-visual interconnection between local and cross-regional restaurant private rooms.

[0129] In the description of this specification, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0130] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0131] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A restaurant system based on real-time audio-visual interconnection via the Internet, characterized in that, This includes interconnected terminals deployed in local restaurants and at least one remote restaurant. Each interconnected terminal includes a visual display device (1), an audio device (2), a visual acquisition device (3), an internet communication module (4), and a dining table (5). The visual display device (1) adopts at least one of LED screen, TV screen, and projection device, with a resolution of not less than 4K UHD (3840×2160 pixels) and a refresh rate of ≥60Hz. It is installed on the wall directly in front of the dining table (5) or embedded in the interior of the dining table, at a distance of 1.2-2.0 meters from the dining table (5). The audio device (2) includes a ring array microphone (21) and a multi-channel surround sound system (22). The microphone (21) is installed at the center of the dining table (5) and contains 6-8 directional microphone units with a sampling rate ≥48kHz and a signal-to-noise ratio ≥70dB. The sound system (22) is symmetrically arranged on both sides of the visual display device (1) with an output power of 80-120W and a frequency response range of 20Hz-20kHz. The visual acquisition device (3) is a high-definition camera with a resolution of ≥1080p, a frame rate of ≥30fps, and a wide-angle lens with a field of view of ≥120°. It is located at any position in the ceiling directly above or on the side bracket of the visual display device (1), at a height of 2.0-2.5 meters from the surface of the dining table (5). The Internet communication module (4) is based on the TCP / IP protocol stack, has built-in WebRTC real-time transmission protocol, supports H.265 video encoding and Opus audio encoding, has a transmission bandwidth of ≥10Mbps / channel, and an end-to-end delay of ≤100ms. Among them, the output end of the visual acquisition device (3) of the local restaurant is connected to the input end of the Internet communication module (4), and the output end of the Internet communication module (4) of the restaurant in another location is connected to the input end of the local visual display device (1) and the audio (22), forming a two-way audio-visual synchronous closed loop.

2. The restaurant system based on real-time audio-visual interconnection via the Internet according to claim 1, characterized in that, The visual display device (1) supports HDR high dynamic range display, has a screen size of 55-85 inches, an adjustable installation angle (tilt angle ±15°), and an anti-glare coating on the screen surface.

3. A restaurant system based on real-time audio-visual interconnection via the Internet according to claim 1, characterized in that, The diameter of the ring array of the microphone (21) is 0.5-0.8 meters, and the angle between adjacent microphone units is 45°-60°. Beamforming technology is used to collect human voices in a directional manner. The speaker (22) is a 5.1 channel system, with the center channel facing the center of the dining table (5).

4. A restaurant system based on real-time audio-visual interconnection via the Internet according to claim 1, characterized in that, The visual acquisition device (3) has a built-in PTZ gimbal mechanism with a horizontal rotation range of ±90° and a vertical rotation range of ±30°. It automatically tracks and locates diners through a face recognition algorithm.

5. A restaurant system based on real-time audio-visual interconnection via the Internet according to claim 1, characterized in that, The Internet communication module (4) includes a main server (41) and an edge computing node (42). The main server (41) is deployed in the cloud and is used to calibrate remote audio and video timestamps. The edge computing node (42) is deployed on the local router for real-time compression of data packets, with a compression rate of 50%-60%.

6. A restaurant system based on real-time audio-visual interconnection via the Internet according to any one of claims 1-5, characterized in that, The positional relationship of each device is as follows: The center point of the visual display device (1) and the center point of the dining table (5) are on the same vertical plane; The center point of the microphone (21) array coincides with the geometric center of the dining table (5), and its height is 0.2-0.3 meters above the tabletop; The optical axis of the visual acquisition device (3) is perpendicular to the surface of the dining table (5), with a downward viewing angle of 10°-20°.

7. A restaurant system based on real-time audio-visual interconnection via the Internet according to claim 6, characterized in that, The main router of the Internet communication module (4) is deployed in the restaurant equipment room and is directly connected to the visual display device (1), audio (22) and visual acquisition device (3) via gigabit network cable. Wireless transmission is only used for communication between the array units of the microphone (21).

8. The mirroring application mode of a restaurant system based on real-time audio-visual interconnection via the Internet, as described in claims 1-7, is characterized in that... The local restaurant is interconnected with a single restaurant in another location, and the local visual display device (1) is divided into two display areas, left and right. The left area displays the complete image captured by the visual acquisition device (3) of the restaurant in another location in real time, while the right area displays local auxiliary information; The audio stream captured by the remote microphone (21) is output through the left channel of the local speaker (22), and a 3-5ms buffer is added during output to align with the screen. The wide-angle lens of the visual acquisition device (3) covers the entire dining area of ​​the different locations, and the image is transmitted without cropping.

9. A supplementary application mode for a restaurant system based on real-time audio-visual interconnection via the Internet, as described in claims 1-7, is characterized in that... The local restaurant is connected to at least two restaurants in other locations at the same time, and the visual display device (1) is divided into a 2×2 grid area; Each grid displays an independent view of a restaurant in another location, with a 16:9 aspect ratio and resolution that adapts to scaling. The audio device (2) uses a dynamic mixing algorithm, and each sound source in a different location is independently mapped to the azimuth angle of a different speaker (22); The Internet communication module (4) enables a multi-point relay server and dynamically allocates bandwidth (single channel ≥ 5Mbps).

10. A single-point application mode of a restaurant system based on real-time audio-visual interconnection via the Internet, as described in claims 1-7, is characterized in that... When only a single restaurant in another location is connected, 85%-90% of the area of ​​the visual display device (1) displays the image from that location. The visual acquisition device (3) is fixed in wide-angle mode and has no PTZ tracking function; The transmission bandwidth of the Internet communication module (4) is reduced to 5-8Mbps, and the latency is ≤120ms; The speaker (22) is simplified to stereo output, and the microphone (21) array unit is reduced to 4.

Citation Information

Cited By

  • 5G MIFI equipment integrating voice and visual identification

    CN121568145A