Display method and display device

By combining a rotatable camera and a sound sensor in the display device, the system identifies the user's position and adjusts the camera angle, solving the problem of the camera being unable to capture images when the user moves, thus improving the user experience of video chat and fitness functions.

CN116097120BActive Publication Date: 2026-05-12HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HISENSE VISUAL TECH CO LTD
Filing Date
2021-05-13
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing display devices, cameras are fixedly installed, resulting in a limited field of view. When users walk out of the camera's shooting area during video chats or workouts, their images cannot be captured, impacting the user experience.

Method used

It adopts a rotatable camera design, combined with a sound collector to collect the sound source information of people, and the controller identifies the angle of the sound source and adjusts the shooting angle of the camera so that the camera always captures the image of the user.

Benefits of technology

It enables the camera to capture images of people even when the user is moving, improving the user experience of functions such as video chat and fitness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116097120B_ABST
    Figure CN116097120B_ABST
Patent Text Reader

Abstract

The application discloses a display method and a display device, wherein a camera can rotate within a preset angle range, a controller is configured to acquire person sound source information collected by a sound collector and perform sound source identification to determine sound source angle information for identifying an azimuth angle of a position where a person is located; based on a current shooting angle of the camera and the sound source angle information, a target rotation direction and a target rotation angle of the camera are determined; and the shooting angle of the camera is adjusted according to the target rotation direction and the target rotation angle, so that a shooting area of the camera is directly opposite to a position where the person is located when the person speaks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 202010848905.X, filed on August 21, 2020, entitled "A Method for Locating and Tracking Audio-Visual Persons"; this application claims priority to Chinese Patent Application No. 202010621070.4, filed on July 1, 2020, entitled "A Method for Adjusting the Shooting Angle of a Camera and a Display Device", the entire contents of which are incorporated herein by reference; this application claims priority to Chinese Patent Application No. 202110014128.3, filed on January 6, 2021, entitled "A Display Device and a Method for Locating and Tracking Audio-Visual Persons", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of television software technology, and more particularly to a display method and display device. Background Technology

[0003] With the rapid development of display devices, their functions will become increasingly richer and their performance more powerful. For example, display devices can realize functions such as web search, IPTV, BBTV, video-on-demand (VOD), digital music, online news, and internet video calls. When using display devices to implement internet video calling, a camera needs to be installed on the display device to capture the user's image. Summary of the Invention

[0004] This application provides a display device, including:

[0005] A camera configured to capture human images and rotate within a preset angle range;

[0006] A sound acquisition device, configured to acquire human voice source information, which refers to the sound information generated when a person interacts with a display device through voice.

[0007] A controller connected to the camera and the sound collector, the controller being configured to: acquire human voice source information collected by the sound collector and the current shooting angle of the camera;

[0008] The voice source information of the person is used to identify the voice source and determine the voice source angle information. The voice source angle information is used to represent the azimuth angle of the person's position when speaking.

[0009] Based on the current shooting angle and sound source angle information of the camera, the target rotation direction and target rotation angle of the camera are determined;

[0010] Adjust the camera's shooting angle according to the target's rotation direction and angle, so that the camera's shooting area is directly facing the position where the person is speaking. Attached Figure Description

[0011] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 The diagram illustrates an operational scenario between a display device and a control device according to some embodiments;

[0013] Figure 2 The diagram illustrates, by way of example, a hardware configuration block diagram of a display device 200 according to some embodiments;

[0014] Figure 3 The diagram illustrates, by way of example, a hardware configuration block diagram of a control device 100 according to some embodiments;

[0015] Figure 4 The diagram illustrates, by way of example, a software configuration schematic of a display device 200 according to some embodiments;

[0016] Figure 5 The image above exemplarily illustrates a schematic diagram of an icon control interface display for an application in a display device 200 according to some embodiments;

[0017] Figure 6 The diagram illustrates a structural block diagram of a display device according to some embodiments;

[0018] Figure 7 The diagram illustrates, by way of example, a preset angle range for camera rotation according to some embodiments;

[0019] Figure 8 The image above exemplifies a scene diagram showing a camera rotating within a preset angle range according to some embodiments;

[0020] Figure 9 The diagram illustrates, by way of example, the range of sound source angles according to some embodiments;

[0021] Figure 10 The flowchart of a method for adjusting the shooting angle of a camera according to some embodiments is illustrated in the example;

[0022] Figure 11 The flowchart of a comparison method for wake-up text according to some embodiments is illustrated in the figure;

[0023] Figure 12The flowchart of a method for sound source identification of human voice source information according to some embodiments is illustrated in the example;

[0024] Figure 13 The flowchart illustrates a method for determining the target rotation direction and target rotation angle of a camera according to some embodiments;

[0025] Figure 14 The image above exemplifies a scenario of adjusting the camera's shooting angle according to some embodiments;

[0026] Figure 15a The diagram illustrates another scenario of adjusting the camera's shooting angle according to some embodiments;

[0027] Figure 15b The image above exemplifies a scene diagram showing the location of a character during speech according to some embodiments;

[0028] Figure 16 This is a schematic diagram of the arrangement of the display device and camera in an embodiment of this application;

[0029] Figure 17 This is a schematic diagram of the camera structure in an embodiment of this application;

[0030] Figure 18a This is a schematic diagram of the display device scene before adjustment in the embodiments of this application;

[0031] Figure 18b This is a schematic diagram of the adjusted display device scenario in the embodiments of this application;

[0032] Figure 19 This is a schematic diagram of a sound source localization scenario in an embodiment of this application;

[0033] Figure 20 This is a schematic diagram of key points in the embodiments of this application;

[0034] Figure 21 This is a schematic diagram of the human portrait center and the image center in an embodiment of this application;

[0035] Figure 22 This is a geometric diagram illustrating the process of calculating the rotation angle in an embodiment of this application.

[0036] Figure 23a This is a schematic diagram of the initial state during the adjustment of the rotation angle in an embodiment of this application;

[0037] Figure 23b This is a schematic diagram showing the result of adjusting the rotation angle in an embodiment of this application;

[0038] Figure 24a This is a schematic diagram of the squatting posture in an embodiment of this application;

[0039] Figure 24b This is a schematic diagram of the standing posture in the embodiments of this application;

[0040] Figure 25a This is a schematic diagram illustrating the initial state display effect of the virtual avatar in the embodiments of this application;

[0041] Figure 25b This is a schematic diagram illustrating the adjusted display effect of the virtual avatar in the embodiments of this application. Detailed Implementation

[0042] To make the objectives, implementation methods and advantages of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only some embodiments of this application, and not all embodiments.

[0043] Based on the exemplary embodiments described in this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the appended claims. Furthermore, although the disclosures in this application are presented by way of one or more exemplary examples, it should be understood that each aspect of these disclosures can also constitute a complete implementation on its own.

[0044] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0045] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities and do not necessarily imply a specific order or sequence, unless otherwise indicated. It should be understood that such terms can be used interchangeably where appropriate, for example, to implement the application in a sequence other than those given in the embodiments illustrated or described herein.

[0046] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclusively include, for example, a product or device that includes a series of components is not necessarily limited to those that are explicitly listed, but may include other components that are not explicitly listed or that are inherent to such product or device.

[0047] As used in this application, the term "module" means any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code capable of performing the functions associated with that element.

[0048] As used in this application, the term "remote control" refers to a component of an electronic device (such as the display device disclosed in this application) that typically allows for wireless control of the electronic device over a short distance. It generally uses infrared and / or radio frequency (RF) signals and / or Bluetooth to connect to the electronic device, and may also include functional modules such as WiFi, wireless USB, Bluetooth, and motion sensors. For example, a handheld touch remote control replaces most of the physical built-in hard buttons in a typical remote control device with a user interface on a touchscreen.

[0049] As used in this application, the term "gesture" refers to user behavior in which a user expresses an expected idea, action, purpose, and / or result through a change in hand shape or hand movement.

[0050] Figure 1 This is a schematic diagram illustrating the operational scenario between the display device and the control unit according to the embodiment. Figure 1 As shown, the user can operate the display device 200 through the smart device 300 or the control device 100.

[0051] The control device 100 can be a remote control. Communication between the remote control and the display device includes infrared protocol communication, Bluetooth protocol communication, and other short-range communication methods, controlling the display device 200 wirelessly or via wired means. Users can control the display device 200 by inputting user commands through buttons on the remote control, voice input, control panel input, etc.

[0052] In some embodiments, a smart device 300 (such as a mobile terminal, tablet computer, computer, laptop computer, etc.) may also be used to control the display device 200. For example, an application running on the smart device may be used to control the display device 200.

[0053] In some embodiments, the display device 200 can also be controlled in ways other than the control device 100 and the smart device 300. For example, it can be controlled by directly receiving the user's voice commands through a module configured inside the display device 200 for acquiring voice commands, or it can be controlled by receiving the user's voice commands through a voice control device set outside the display device 200.

[0054] In some embodiments, the display device 200 also communicates with the server 400. The display device 200 may communicate via a local area network (LAN), wireless local area network (WLAN), and other networks. The server 400 may provide the display device 200 with various content and interactive features.

[0055] Figure 2 An exemplary block diagram of the configuration of the control device 100 according to an exemplary embodiment is shown. Figure 3As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, a memory 190, and a power supply 180. The control device 100 can receive user input operation commands and convert the operation commands into commands that the display device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the display device 200.

[0056] Figure 2 A hardware configuration block diagram of a display device 200 according to an exemplary embodiment is shown.

[0057] Display device 200 includes at least one of tuner / demodulator 210, communicator 220, detector 230, external device interface 240, controller 250, display 275, audio output interface 285, memory 260, power supply 290, and user interface 265.

[0058] The display 275 includes a display screen assembly for presenting images, a driving assembly for driving image display, a component for receiving image signals from the controller output, and a user interface for displaying video content, image content, menu control interface, and user control UI.

[0059] The display 275 may be a liquid crystal display, an OLED display, or a projection display, and may also be a projection device and a projection screen.

[0060] The communicator 220 is a component used to communicate with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of the following: a Wi-Fi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The display device 200 can establish the transmission and reception of control signals and data signals with the external control device 100 or the server 400 through the communicator 220.

[0061] The user interface can be used to receive control signals from the control device 100 (such as an infrared remote control).

[0062] Detector 230 is used to collect signals from the external environment or to interact with the external environment. For example, detector 230 includes a light receiver, a sensor for collecting ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0063] The external device interface 240 may include, but is not limited to, one or more of the following: High Definition Multimedia Interface (HDMI), analog or high-definition component input interface (component), composite video input interface (CVBS), USB input interface (USB), RGB port, etc. It may also be a composite input / output interface formed by multiple interfaces mentioned above.

[0064] The controller 250 and the tuner 210 can be located in different separate devices, that is, the tuner 210 can also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0065] The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory 260. The controller 250 controls the overall operation of the display device 200. For example, in response to receiving a user command to select a UI object to display on the monitor 275, the controller 250 can perform operations related to the object selected by the user command.

[0066] The object can be any of the optional objects, such as hyperlinks, icons, or other operable controls. Operations related to the selected object include: displaying links to hyperlinked pages, documents, images, etc., or performing operations corresponding to the program associated with the icon.

[0067] In some embodiments, the user can input user commands through a graphical user interface (GUI) displayed on the display 275, and the user input interface receives the user input commands through the GUI. Alternatively, the user can input user commands by inputting specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors.

[0068] A "user interface" refers to the medium through which an application or operating system interacts and exchanges information with the user. It converts information from its internal form to a form that the user can accept. A common form of user interface is the graphical user interface (GUI), which is a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0069] See Figure 4In some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the Android runtime and system library layer (referred to as the "System Runtime Layer"), and the kernel layer.

[0070] In some embodiments, at least one application runs in the application layer. These applications may be built-in Windows programs, system settings programs, clock programs, camera applications, etc., or applications developed by third-party developers, such as HiSee programs, karaoke programs, magic mirror programs, etc. In specific implementations, the application packages in the application layer are not limited to the examples above, and may actually include other application packages. This application embodiment does not impose any limitations on this.

[0071] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications in the application layer. Applications can access system resources and obtain system services during execution through the API interface.

[0072] like Figure 4 As shown, the application framework layer in this embodiment includes managers, content providers, etc., wherein the managers include at least one of the following modules: ActivityManager, which interacts with all activities running in the system; LocationManager, which provides access to system location services for system services or applications; PackageManager, which retrieves various information related to application packages currently installed on the device; NotificationManager, which controls the display and clearing of notification messages; and WindowManager, which manages icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0073] In some embodiments, the Activity Manager is used to: manage the lifecycle of each application and general navigation back functions, such as controlling the exit of an application (including switching the currently displayed user interface in the display window to the system desktop), opening an application, and going back (including switching the currently displayed user interface in the display window to the previous level user interface).

[0074] In some embodiments, the window manager is used to manage all window programs, such as obtaining the screen size, determining whether there is a status bar, locking the screen, capturing the screen, and controlling changes to the display window (e.g., shrinking the display window, shaking the display, distorting the display, etc.).

[0075] In some embodiments, the system runtime library layer provides support for the upper layer, namely the framework layer. When the framework layer is used, the Android operating system runs the C / C++ libraries contained in the system runtime library layer to implement the functions that the framework layer needs to perform.

[0076] In some embodiments, the kernel layer is a layer between hardware and software. For example... Figure 4 As shown, the kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, touch sensor, pressure sensor, etc.).

[0077] In some embodiments, the kernel layer also includes a power driver module for power management.

[0078] In some embodiments, Figure 4 The software programs and / or modules corresponding to the software architecture in the document are stored in [the relevant database]. Figure 2 or Figure 3 In the first or second memory shown.

[0079] In some embodiments, taking the Magic Mirror application (photo-taking application) as an example, when the remote control receiver receives a remote control input operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the input operation into a raw input event (including the value of the input operation, the timestamp of the input operation, etc.). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer, identifies the control corresponding to the input event based on the current focus position, and determines whether the input operation is a confirmation operation. The control corresponding to the confirmation operation is the Magic Mirror application icon. The Magic Mirror application calls the interface of the application framework layer to start the Magic Mirror application, and then calls the kernel layer to start the camera driver, thereby capturing still images or videos through the camera.

[0080] In some embodiments, for touch-enabled display devices, taking split-screen operation as an example, the display device receives user input operations (such as split-screen operation) on the display screen. The kernel layer can generate corresponding input events based on the input operations and report the events to the application framework layer. The activity manager of the application framework layer sets the window mode (such as multi-window mode), window position, and size corresponding to the input operation. The window management of the application framework layer draws the window according to the settings of the activity manager, and then sends the drawn window data to the display driver of the kernel layer. The display driver then displays the corresponding application interface in different display areas of the screen.

[0081] In some embodiments, such as Figure 5 As shown, the application layer contains at least one application that can display corresponding icon controls on the display, such as: live TV application icon control, video-on-demand application icon control, media center application icon control, application center icon control, game application icon control, etc.

[0082] In some embodiments, a live TV application can provide live TV from different signal sources. For example, the live TV application can provide a TV signal using input from cable television, terrestrial broadcasting, satellite services, or other types of live TV services. Furthermore, the live TV application can display the video of the live TV signal on display device 200.

[0083] In some embodiments, a video-on-demand application may provide video from different storage sources. Unlike live TV applications, video-on-demand provides video display from certain storage sources. For example, video-on-demand may come from a cloud storage server or from local hard drive storage containing existing video programs.

[0084] In some embodiments, a media center application may be an application that provides playback of various multimedia content. For example, a media center may provide services that, unlike live TV or video-on-demand, allow users to access various images or audio through the media center application.

[0085] In some embodiments, the application center may provide a storage for various applications. An application may be a game, an application, or other applications related to a computer system or other device but capable of running on a smart TV. The application center may obtain these applications from various sources, store them in local storage, and then make them runnable on the display device 200.

[0086] First aspect:

[0087] In some embodiments, applications utilizing the camera in the display device include "HiSee," "Mirror," "Youxuemao," and "Fitness," enabling functions such as "video chat," "watch and chat simultaneously," and "fitness." "HiSee" is a video chat application that allows one-click chat between mobile phones and TVs, and between TVs. "Mirror" is an application that provides a mirror service; by opening the camera through the Mirror application, users can use their smart TV as a mirror. "Youxuemao" is an application that provides learning functions. When implementing the "watch and chat simultaneously" function, users can watch video programs while simultaneously engaging in a video call using the "HiSee" application. The "Fitness" function can simultaneously display fitness guidance videos and images captured by the camera showing the user performing the corresponding movements according to the fitness guidance videos on the display device's screen, allowing users to check in real time whether their movements are correct.

[0088] Since users may not remain stationary when using display devices for "video chat," "watch and chat," or "exercise," they can perform these functions while moving around. However, in existing display devices, the camera is fixedly mounted, with its viewing angle perpendicular to the monitor. Furthermore, the camera's field of view is limited, typically between 60° and 75°. This means the camera's shooting area is defined by a 60°–75° angle extending simultaneously to the left and right of the camera's viewing angle.

[0089] If a user steps out of the camera's field of view, the camera will not capture an image containing the user's image, resulting in the image not being displayed on the monitor. In a video chat scenario, the other user will not be able to see the user; in a fitness scenario, the monitor will not display an image of the user performing fitness exercises, making it impossible for the user to see their own movements and determine if they are performed correctly, thus impacting the user experience.

[0090] Figure 6 The diagram illustrates a structural block diagram of a display device according to some embodiments. To ensure the camera can still capture images of the user even when the user has left the camera's shooting area, see [link to relevant documentation]. Figure 6 This application provides a display device including a camera 232, a sound acquisition unit 231, and a controller 250. The camera is used to capture human images. Instead of being fixedly installed, the camera is rotatably mounted on the display device. Specifically, the camera 232 is rotatably mounted on the top of the display and can rotate along the top of the display.

[0091] Figure 7 The diagram illustrates, by way of example, a preset angle range for camera rotation according to some embodiments; Figure 8The image illustrates a scene where a camera rotates within a preset angle range according to some embodiments. See also... Figure 7 and Figure 8 The preset camera 232 can rotate within a preset angle range, and rotates horizontally. In some embodiments, the preset angle range is 0° to 120°, that is, when facing the display, the left side of the user is 0° and the right side of the user is 120°. Taking the state where the center line of the camera 232's viewing angle is perpendicular to the display as the initial state, the camera can rotate 60° to the left from the initial state and 60° to the right from the initial state; the position where the center line of the camera's viewing angle is perpendicular to the display is the 60° position of the camera.

[0092] The display device provided in this application embodiment enables the camera to rotate by triggering sound source information, automatically identifying the user's real-time location and adjusting the camera's shooting angle so that the camera can always capture images containing human figures. Therefore, in some embodiments, the display device uses a sound acquisition unit 231 to collect human sound source information.

[0093] To ensure the accuracy of sound source acquisition, multiple sound acquisition units can be set in the display device. In some embodiments, four sound acquisition units 231 are set in the display device, and the four sound acquisition units 231 can be arranged in a linear positional relationship. In some embodiments, the sound acquisition units can be microphones, and the four microphones are arranged linearly to form a microphone array. During sound acquisition, the four sound acquisition units 231 receive the sound information generated when the same user interacts with the display device through voice.

[0094] Figure 9 The diagram illustrates, by way of example, a sound source angle range according to some embodiments. When a user speaks, the sound is received 360°. Therefore, when the user is in front of the display device, the sound source angle range is 0° to 180°. Similarly, when the user is behind the display device, the sound source angle range is also 0° to 180°. See also... Figure 9 Taking the user's position facing the display device as an example, the user is at 0° horizontally when they are to the left of the sound acquisition device, and at 180° horizontally when they are to the right of the sound acquisition device.

[0095] See you again Figure 7 and Figure 9 The 30° position of the sound source is equal to the 0° position of the camera, the 90° position of the sound source is equal to the 60° position of the camera, and the 150° position of the sound source is equal to the 120° position of the camera.

[0096] The controller 250 is connected to both the camera 232 and the sound acquisition unit 231. The controller receives and identifies the voice source information collected by the sound acquisition unit, determines the location angle of the person, and then determines the angle at which the camera needs to rotate. The controller adjusts the camera's shooting angle according to the determined angle, ensuring that the camera's shooting area is directly facing the person's location when they are speaking, thus achieving the goal of adjusting the camera's shooting angle based on the person's position to capture an image containing the person.

[0097] Figure 10 The flowchart illustrates an exemplary method for adjusting the camera shooting angle according to some embodiments. An embodiment of this application provides a display device in which, when adjusting the camera shooting angle based on the position of a person, the controller is configured to perform... Figure 10 The methods for adjusting the camera's shooting angle shown include:

[0098] S1. Obtain the voice source information of the person collected by the sound collector and the current shooting angle of the camera.

[0099] In some embodiments, when the controller in the display device drives the camera to rotate to adjust the camera's shooting angle, it needs to determine the source information of the person's voice when the person interacts with the display device through voice. The source information of the person's voice refers to the sound information generated when the person interacts with the display device through voice.

[0100] Voice source information can determine the location and angle of a person speaking. However, to accurately determine the angle the camera needs to be adjusted, the current state of the camera, i.e., the current shooting angle, must first be obtained. The current shooting angle of the camera can only be obtained when the camera is stationary to ensure its accuracy, and thus the accuracy of determining the angle that needs to be adjusted.

[0101] Therefore, before acquiring the current shooting angle of the camera, the controller is further configured to perform the following steps:

[0102] Step 11: Check the current operating status of the camera.

[0103] Step 12: If the camera is currently in a rotating state, wait for the camera to finish rotating.

[0104] Step 13: If the camera is currently in a non-rotating state, obtain the current shooting angle of the camera.

[0105] The controller is equipped with a motor control service, which is used to drive the camera to rotate, acquire the camera's operating status, and the camera's orientation angle.

[0106] The motor control service monitors the camera's operating status in real time. The controller queries the camera's current operating status by calling the motor control service. The camera's current operating status can indicate the camera's orientation angle and whether the camera is rotating.

[0107] If the camera is rotating, its current shooting angle cannot be obtained, otherwise an accurate value cannot be determined. Therefore, when the camera is rotating, you must wait for the camera to complete the rotation after executing the previous instruction, and then execute the step of obtaining the current shooting angle after the camera has stopped.

[0108] If the camera is not rotating, i.e., the camera is stationary, then the step of obtaining the current shooting angle of the camera can be performed.

[0109] S2. Perform sound source identification on the voice source information of the person and determine the sound source angle information. The sound source angle information is used to represent the orientation angle of the person's position when speaking.

[0110] After acquiring the voice source information generated by the interaction between the person and the display device, the controller needs to perform sound source recognition to determine the person's position when speaking, specifically the orientation angle, that is, whether the person is located to the left, right or directly facing the sound acquisition device, and then adjust the camera's shooting angle according to the person's position.

[0111] When a person interacts with a display device, such as in a video call, their voice may be conversing with the other user while they remain within the camera's field of view. If the controller attempts to adjust the camera's angle at this time, the operation will be ineffective.

[0112] Therefore, in order to accurately determine whether the camera's shooting angle needs to be adjusted based on the person's voice source information, it is necessary to first analyze the person's voice source information and determine whether the person's voice source information is the information that triggers the camera adjustment.

[0113] In some embodiments, wake-up text for triggering camera angle adjustment can be pre-stored in the controller. For example, "Hisense Xiaoju" can be customized as the wake-up text for sound source recognition. A person triggers the camera angle adjustment process by recognizing the voice as "Hisense Xiaoju". The wake-up text can also be customized to other words; this embodiment does not impose specific limitations.

[0114] Figure 11 The diagram illustrates a flowchart of a method for comparing wake-up text according to some embodiments. Specifically, see [link to relevant documentation]. Figure 11Before performing sound source identification and determining the sound source angle information, the controller is further configured to perform the following steps:

[0115] S021. Extract text from the voice source information of the characters to obtain the voice interaction text.

[0116] S022. Compare the voice interaction text and the preset wake-up text. The preset wake-up text refers to the text used to trigger the sound source recognition process.

[0117] S023. If the voice interaction text matches the preset wake-up text, then perform the step of voice source recognition for the voice source information.

[0118] In some embodiments, after acquiring the voice source information, the controller first performs text extraction to extract the voice interaction text when the person interacts with the display device via voice. The extracted voice interaction text is compared with the preset wake-up text. If the comparison is inconsistent, for example, the person's voice is not "Hisense Xiaoju" but other interactive content, it means that the current person's voice is not the voice that triggers the adjustment of the camera shooting angle, and the controller does not need to execute the relevant steps for adjusting the camera shooting angle.

[0119] If the comparison is consistent, it means that the current person's voice is the voice that triggers the adjustment of the camera shooting angle. For example, if the person's voice is the pre-set "Hisense Xiaoju", the controller can continue to execute the subsequent steps of adjusting the camera shooting angle.

[0120] When the controller determines that the voice source is a wake-up voice, i.e., the trigger voice for adjusting the camera's shooting angle, it needs to execute the subsequent sound source recognition process.

[0121] Since the display device is equipped with multiple sound acquisition units, these units can collect multiple sets of sound source information of the same person when they speak. Therefore, when the controller obtains the sound source information of the person collected by the sound acquisition units, it can obtain the sound source information of the person generated when they speak, that is, the controller will obtain multiple sets of sound source information.

[0122] Figure 12 The document exemplarily illustrates a flowchart of a method for sound source identification of a person's voice source information according to some embodiments. When multiple sets of sound collectors collect the same wake-up text, since the distance between each sound collector and the person is not the same, the sound source information of each person can be identified to determine the directional angle of the person's speech, i.e., the sound source angle information. Specifically, see [link to documentation]. Figure 12 The controller, after performing sound source identification on the voice source information and determining the sound source angle information, is further configured to perform the following steps:

[0123] S21. Perform sound source identification for each person's sound source information, and calculate the speech time difference generated when multiple sound acquisition devices collect the corresponding person's sound source information.

[0124] S22. Based on the speech time difference, calculate the sound source angle information of the person's location when speaking.

[0125] Each sound acquisition device has the same frequency response and its sampling clock is synchronized. However, since the distance between each sound acquisition device and the person is not the same, the time when each sound acquisition device can acquire speech is not the same, and there will be a time difference in acquisition between multiple sound acquisition devices.

[0126] In some embodiments, the angle and distance between the sound source and the array can be calculated using a sound acquisition array to track the sound source at the location of a person speaking. Based on TDOA (Time Difference of Arrival) sound source localization technology, the time difference between the arrival of the signal at each microphone is estimated to obtain a set of equations for the sound source's position coordinates. Solving the set of equations then yields the precise azimuth coordinates of the sound source, i.e., the sound source angle information.

[0127] In some embodiments, in step S21, the controller, after performing sound source identification for each of the human voice source information and calculating the speech time difference generated when the multiple sets of sound collectors collect the corresponding human voice source information, is further configured to perform the following steps:

[0128] Step 211: Extract environmental noise, the sound source signal of the person's speech, and the propagation time of the person's speech to each sound acquisition device from the voice source information.

[0129] Step 212: Determine the received signal of each sound collector based on the ambient noise, sound source signal, and propagation time.

[0130] Step 213: Using the cross-correlation time delay estimation algorithm, process the received signal of each sound collector to obtain the speech time difference generated by each pair of sound collectors when collecting the corresponding voice source information.

[0131] When calculating the speech time difference between two sound acquisition devices, the direction of arrival (DOA) estimation can be achieved using the sound acquisition device array. The DOA estimation algorithm is used to calculate the time difference between the arrival times of the sound in different sound acquisition device arrays.

[0132] In a sound source localization system, the target signal received by each element of the sound acquisition array originates from the same sound source. Therefore, there is a strong correlation between the signals from different channels. By calculating the correlation function between any two signals, the time delay between any two observed signals from the sound acquisition devices can be determined, i.e., the speech time difference.

[0133] The voice source information generated by a person during speech includes environmental noise and the voice source signal of the person speaking. Furthermore, the propagation time of the person's voice to each sound acquisition device can be extracted from the voice source information, and the received signal of each sound acquisition device can be calculated.

[0134] x i (t)=α i s(t-τ i )+n i (t);

[0135] In the formula, x i τ is the received signal of the i-th sound collector, s(t) is the sound source signal when the person speaks, and τ is the sound source signal when the person speaks. i Let n be the propagation time of a person's voice to the i-th sound collector. i (t) represents environmental noise, α i This is a correction factor.

[0136] The received signals from each sound acquisition unit are processed using a cross-correlation time delay estimation algorithm to estimate the time delay, as shown below: In the formula, The time delay between the i-th and (i+1)-th sound collectors is the speech time difference.

[0137] Substituting the received signal model of each sound acquisition device, we get:

[0138]

[0139] Since s(t) and n i Since (t) are uncorrelated, the above equation can be simplified to:

[0140]

[0141] in, n i With n i+1 If the noise is uncorrelated Gaussian white noise, the above equation further simplifies to:

[0142]

[0143] From the properties of the cross-correlation time delay estimation algorithm, it can be seen that when hour, The maximum value represents the time delay between the two sound acquisition devices, i.e., the speech time difference.

[0144] In the actual signal processing model of the sound acquisition array, due to the effects of reverberation and noise, The peak value is not obvious, which reduces the accuracy of latency estimation. To sharpen... The peak value can be obtained by weighting the cross-power spectrum in the frequency domain based on prior knowledge of the signal and noise, thereby suppressing noise and reverberation interference. Finally, an inverse Fourier transform is performed to obtain the generalized cross-correlation function.

[0145] in This represents the frequency domain weighting function.

[0146] Finally, PHAT weighting is used to smooth the interaction rate spectrum between signals, resulting in the final speech time difference generated by each pair of sound acquisition devices when acquiring the corresponding human voice source information. The PHAT-weighted cross-power spectrum approximates the expression for the unit impulse response, highlighting the peak value of the time delay. This effectively suppresses reverberation noise and improves the accuracy and precision of time delay (speech time difference) estimation.

[0147] In some embodiments, in step S22, the controller, when performing the calculation of the sound source angle information of the person's location during speech based on the speech time difference, is further configured to perform the following steps:

[0148] Step 221: Obtain the speed of sound in the current environment, the coordinates of each sound collector, and the number of sound collectors set.

[0149] Step 222: Determine the number of combination pairs of sound collectors based on the number of sound collectors set. The number of combination pairs refers to the number of combinations obtained by combining two sound collectors in pairs.

[0150] Step 223: Based on the speech time difference, sound velocity, and coordinates of each pair of sound acquisition devices, establish a set of vector relationship equations. The number of vector relationship equations is the same as the number of combination pairs.

[0151] Step 224: Solve the system of vector relation equations to obtain the vector value of the propagation vector of the unit plane wave of the sound source at the location where the person is speaking.

[0152] Step 225: Calculate the sound source angle information of the person's position when speaking, based on the vector value.

[0153] After calculating the speech time difference between two sound collectors according to the method provided in the foregoing embodiments, the sound source angle information of the person's position when speaking can be calculated based on each speech time difference.

[0154] When calculating the sound source angle information, multiple sets of vector relationship equations need to be established. To ensure the accuracy of the calculation results, the number of equation sets can be set to be the same as the number of combinations obtained by pairwise combinations of sound collectors. Therefore, if the number of sound collectors N is set, then there are N(N-1) / 2 pairs of combinations between all sound collectors.

[0155] When establishing the system of vector equations, obtain the speed of sound c in the current environment and the coordinates of each sound collector. Let the coordinates of the k-th sound collector be (x... k ,y k ,z k Meanwhile, the propagation vector of the sound source unit plane wave at the location where the character is speaking is set as u = (u, v, w). By solving for the vector value of the propagation vector of the sound source unit plane wave at the location where the character is speaking, the angle information of the sound source can be determined.

[0156] Based on the speech time difference between each pair of sound collectors Speed ​​of sound c, coordinates (x, y) of each sound collector k ,y k ,z k Let the propagation vector of the unit plane wave at the location of the sound source during the character's speech be (u, v, w). Establish a system of N(N-1) / 2 vector relationship equations:

[0157] This equation represents the set of vector relationship equations established between the i-th and j-th sound collectors.

[0158] Taking N=3 as an example, the following system of equations can be established:

[0159] (The system of vector relationship equations established between the first and second sound collectors);

[0160]

[0161] (The system of vector relationship equations established between the first and third sound collectors);

[0162]

[0163] (The system of vector relationship equations established between the third sound collector and the second sound collector).

[0164] Rewrite the above three vector relationship equations in matrix form:

[0165]

[0166] Solving for u = (u, v, w) from the above matrix, and then using the sine and cosine relationship, the angle value can be obtained:

[0167] That is, the location and angle of the sound source when a person is speaking.

[0168] S3. Based on the current shooting angle of the camera and the angle of the sound source, determine the target rotation direction and target rotation angle of the camera.

[0169] The controller identifies the sound source by analyzing the voice information to determine the sound source angle information, which represents the location of the person speaking. The sound source angle information identifies the person's current position, and the camera's current shooting angle identifies the camera's current position. Based on the angle difference between the two positions, the target rotation angle that the camera needs to rotate, as well as the target rotation direction of the camera, can be determined.

[0170] Figure 13 The document exemplarily illustrates a method flowchart for determining the target rotation direction and target rotation angle of a camera according to some embodiments. Specifically, see [link to relevant documentation]. Figure 13 The controller, based on the camera's current shooting angle and sound source angle information, determines the camera's target rotation direction and target rotation angle, and is further configured to perform the following steps:

[0171] S31. Convert the sound source angle information into the camera's coordinate angle.

[0172] Since the sound source angle information represents the location angle of a person, in order to accurately calculate the azimuth angle that the camera needs to adjust based on the sound source angle information and the current shooting angle of the camera, the sound source angle information of the person can be converted into the coordinate angle of the camera, that is, the coordinate angle of the camera can be used to replace the sound source angle information of the person.

[0173] Specifically, the controller, in converting the sound source angle information into camera coordinate angles, is further configured to perform the following steps:

[0174] Step 311: Obtain the range of the sound source angle when the person is speaking and the preset range of the camera rotation angle.

[0175] Step 312: Calculate the angle difference between the sound source angle range and the preset angle range, and use half of the angle difference as the conversion angle.

[0176] Step 313: Calculate the angle difference between the angle corresponding to the sound source angle information and the converted angle, and use the angle difference as the coordinate angle of the camera.

[0177] Since the sound source angle range and the camera's preset angle range are different (0°–120°, 0°–180°), the camera's coordinate angle cannot be directly used to replace the sound source angle information. Therefore, the angle difference between the sound source angle range and the preset angle range is calculated first, and then half of the angle difference is calculated. This half-value is used as the conversion angle when converting the sound source angle information to the camera's coordinate angle.

[0178] The angle difference between the sound source angle range and the preset angle range is 60°, and half of this difference is 30°. 30° is used as the conversion angle. Finally, the angle difference between the angle corresponding to the sound source angle information and the conversion angle is calculated; this is the coordinate angle of the camera converted from the sound source angle information.

[0179] For example, if the person is located to the left of the sound collector, the controller determines the angle of the sound source by acquiring the sound source information of the person from multiple sound collectors. The angle corresponding to this angle is 50°, and the angle conversion is 30°. Therefore, the angle difference is calculated to be 20°, which means that the 50° corresponding to the sound source angle information is replaced with the coordinate angle of the camera, which is 20°.

[0180] If the person is located to the right of the sound collector, the controller determines the angle of the sound source by acquiring the sound source information of the person from multiple sound collectors. The angle corresponding to this angle is 130°, and the conversion angle is 30°. Therefore, the angle difference is calculated to be 100°, which means that the 130° corresponding to the sound source angle information is replaced with the coordinate angle of the camera, which is 100°.

[0181] S32. Calculate the angle difference between the camera's coordinate angle and the camera's current shooting angle, and use the angle difference as the target rotation angle of the camera.

[0182] The camera's coordinate angle is used to identify the angle of the person's position within the camera's coordinate system. Therefore, based on the angle difference between the camera's current shooting angle and the camera's coordinate angle, the target rotation angle that the camera needs to rotate can be determined.

[0183] For example, if the current shooting angle of the camera is 100° and the coordinate angle of the camera is 20°, it means that the current shooting area of ​​the camera is not aligned with the position of the person. The two are 80° apart. Therefore, the camera needs to be rotated 80° before the shooting area of ​​the camera can be aligned with the position of the person. That is, the target rotation angle of the camera is 80°.

[0184] S33. Determine the target rotation direction of the camera based on the angle difference.

[0185] Since the camera's position is determined by facing the display device, with the left side as the 0° position and the right side as the 120° position, after determining the angle difference based on the camera's coordinate angle and current shooting angle, if the current shooting angle is greater than the coordinate angle, it means the camera's shooting angle is to the right of the person's position, and the angle difference is negative; if the current shooting angle is less than the coordinate angle, it means the camera's shooting angle is to the left of the person's position, and the angle difference is positive.

[0186] In some embodiments, the target rotation direction of the camera can be determined based on the sign of the angle difference. If the angle difference is positive, it means that the camera's shooting angle is to the left of the person's position. In this case, in order for the camera to capture an image of the person, the camera's shooting angle needs to be adjusted to the right, thus determining the target rotation direction of the camera as rightward rotation.

[0187] If the angle difference is negative, it means that the camera's shooting angle is to the right of the person's position. In this case, in order for the camera to capture an image of the person, the camera's shooting angle needs to be adjusted to the left, thus determining that the camera's target rotation direction is to the left.

[0188] For example, Figure 14 The image illustrates a scenario of adjusting the camera's shooting angle according to some embodiments. See also... Figure 14 If the angle corresponding to the sound source information of the person is 50°, then the converted camera coordinate angle is 20°; the current shooting angle of the camera is 100°, that is, the center line of the camera's field of view is located to the right of the person's position, and the calculated angle difference is -80°. It can be seen that the angle difference is negative, in which case the camera needs to be adjusted to rotate 80° to the left.

[0189] Figure 15a The diagram illustrates another scenario of adjusting the camera's shooting angle according to some embodiments. See also... Figure 15a If the angle corresponding to the sound source of the person is 120°, then the converted camera coordinate angle is 90°; the current shooting angle of the camera is 40°, that is, the center line of the camera's field of view is located to the left of the person's position, and the calculated angle difference is 50°. It can be seen that the angle difference is a positive value. At this time, the camera needs to be adjusted to rotate 50° to the right.

[0190] S4. Adjust the camera's shooting angle according to the target's rotation direction and angle, so that the camera's shooting area is directly facing the position where the person is speaking.

[0191] Once the controller determines the target rotation direction and angle required for the camera to adjust its shooting angle, it can adjust the camera's shooting angle accordingly, positioning the camera's shooting area directly in front of the person's location. This allows the camera to capture images including the person, thus enabling the camera to adjust its shooting angle based on the person's position.

[0192] Figure 15b The diagram illustrates, exemplarily, a scene depicting the location of a person speaking according to some embodiments. Since the preset angle range of the camera differs from the angle range of the sound source during speech, please refer to the angle diagram for details. Figure 15b There is a 30° angle difference between the 0° position of the preset angle range and the 0° position of the sound source angle range. Similarly, there is also a 30° angle difference between the 120° position of the preset angle range and the 180° position of the sound source angle range.

[0193] So, if a person is interacting with a display device and their position happens to be within a 30° angle range, such as... Figure 15b The location of person (a) or person (b) shown in the diagram. At this time, when the controller converts the sound source angle information into the camera coordinate angle in the aforementioned step S31, there will be cases where the camera coordinate angle converted from the sound source angle information of the person is negative, or is greater than the maximum value of the camera's preset angle range, that is, the converted camera coordinate angle is not within the camera's preset angle range.

[0194] For example, if the sound source angle information corresponding to the location of person (a) is 20°, and the conversion angle is 30°, then the calculated camera coordinate angle is -10°. If the sound source angle information corresponding to the location of person (b) is 170°, and the conversion angle is 30°, then the calculated camera coordinate angle is 140°. It can be seen that the camera coordinate angles obtained from the conversions based on the locations of person (a) and person (b) respectively both exceed the camera's preset angle range.

[0195] If the camera's coordinate angles all exceed the camera's preset angle range, it means the camera cannot rotate to the position corresponding to the camera's coordinate angle (the location of the person's voice). Since the camera's visible angle range is between 60° and 75°, this means that when the camera is rotated to the 0° or 120° position, there is a 30° angle difference between the 0° position of the preset angle range and the 0° position of the sound source angle range, and a 30° angle difference between the 120° position of the preset angle range and the 180° position of the sound source angle range.

[0196] Therefore, if the person's position is within a 30° angle difference between the 0° position of the preset angle range and the 0° position of the sound source angle range, or within a 30° angle difference between the 120° position of the preset angle range and the 180° position of the sound source angle range, then in order to capture an image containing the person, the camera's shooting angle is adjusted according to the position corresponding to the minimum or maximum value of the camera's preset angle range.

[0197] In some embodiments, the controller is further configured to perform the following steps: when the sound source angle information of a person is converted into the coordinate angle of the camera and exceeds the preset angle range of the camera, determine the target rotation direction and target rotation angle of the camera based on the angle difference between the current shooting angle of the camera and the minimum or maximum value of the preset angle range.

[0198] For example, if there is an angle difference of 30° between the person (a) and the 0° position within the preset angle range, i.e., the sound source angle corresponding to the person (a)'s sound source angle information is 20°, and the camera's current shooting angle is 50°, the angle difference is calculated based on the minimum value of the camera's preset angle range (0°) and the current shooting angle (50°). If the angle difference is -50°, the target rotation direction of the camera is determined to be leftward, and the target rotation angle is 50°. At this time, the camera's field of view center line (a) coincides with the camera's 0° line.

[0199] If there is a 30° angle difference between the person (b) located at a position within the preset angle range of 120° and the position within the preset angle range of 180°, that is, the sound source angle corresponding to the sound source angle information of person (b) is 170°, and the current shooting angle of the camera is 50°, then the angle difference is calculated based on the maximum value of the preset angle range of the camera (120°) and the current shooting angle (50°). If the angle difference is 70°, then the target rotation direction of the camera is determined to be to the right, and the target rotation angle is 70°. At this time, the center line of the camera's field of view (b) coincides with the camera's 120° line.

[0200] Therefore, even if the angle of the sound source corresponding to the location of the person exceeds the preset angle range when the camera is rotated, the display device provided in this application embodiment can still rotate the camera to the minimum or maximum value corresponding to the preset angle range according to the location of the person, and capture an image containing the person according to the visible angle coverage of the camera.

[0201] As can be seen, the display device provided in this application embodiment includes a camera that can rotate within a preset angle range. The controller is configured to acquire voice source information collected by a sound collector and perform sound source recognition to determine the sound source angle information used to identify the location of the person. Based on the current shooting angle of the camera and the sound source angle information, the controller determines the target rotation direction and target rotation angle of the camera. According to the target rotation direction and target rotation angle, the controller adjusts the shooting angle of the camera so that the camera's shooting area is directly facing the location of the person speaking. Therefore, the display device provided in this application can trigger the rotation of the camera using voice source information, automatically identify the user's real-time location, and adjust the camera's shooting angle so that the camera can always capture an image containing the person's image.

[0202] Figure 10 The diagram illustrates a flowchart of a method for adjusting the camera shooting angle according to some embodiments. See also... Figure 10 This application provides a method for adjusting the shooting angle of a camera, which is executed by the controller in the display device provided in the foregoing embodiments. The method includes:

[0203] S1. Obtain the voice source information of the person collected by the sound collector and the current shooting angle of the camera. The voice source information of the person refers to the sound information generated when the person interacts with the display device through voice.

[0204] S2. Perform sound source identification on the voice source information of the person and determine the sound source angle information. The sound source angle information is used to represent the azimuth angle of the person's position when speaking.

[0205] S3. Based on the current shooting angle and sound source angle information of the camera, determine the target rotation direction and target rotation angle of the camera;

[0206] S4. Adjust the camera's shooting angle according to the target's rotation direction and angle, so that the camera's shooting area is directly facing the position where the person is speaking.

[0207] In some embodiments of this application, before performing sound source identification on the voice source information and determining the sound source angle information, the method further includes: extracting text from the voice source information to obtain voice interaction text; comparing the voice interaction text with a preset wake-up text, wherein the preset wake-up text refers to the text used to trigger the sound source identification process; if the voice interaction text matches the preset wake-up text, then the step of performing sound source identification on the voice source information is performed.

[0208] In some embodiments of this application, multiple sets of sound acquisition devices are included. The controller acquires the human voice source information acquired by the sound acquisition devices by acquiring the human voice source information generated by each sound acquisition device when the human is speaking. The step of performing sound source identification on the human voice source information and determining the sound source angle information includes: performing sound source identification on each human voice source information separately, calculating the speech time difference generated by the multiple sets of sound acquisition devices when acquiring the corresponding human voice source information; and calculating the sound source angle information of the position of the human when speaking based on the speech time difference.

[0209] In some embodiments of this application, the step of performing sound source identification on each of the aforementioned voice source information and calculating the speech time difference generated by multiple sets of sound collectors when collecting the corresponding voice source information includes: extracting environmental noise, the sound source signal of the voice, and the propagation time of the voice to each sound collector from the voice source information; determining the received signal of each sound collector based on the environmental noise, the sound source signal, and the propagation time; and processing the received signal of each sound collector using a cross-correlation delay estimation algorithm to obtain the speech time difference generated by every two sound collectors when collecting the corresponding voice source information.

[0210] In some embodiments of this application, the step of calculating the sound source angle information of the location of the person speaking based on the speech time difference includes: obtaining the sound speed in the current environment, the coordinates of each sound collector, and the number of sound collectors; determining the number of combination pairs of sound collectors based on the number of sound collectors, where the number of combination pairs refers to the number of combinations obtained by combining two sound collectors in pairs; establishing a set of vector relationship equations based on the speech time difference, sound speed, and coordinates of each pair of sound collectors, where the number of vector relationship equations is the same as the number of combination pairs; solving the set of vector relationship equations to obtain the vector value of the sound source unit plane wave propagation vector at the location of the person speaking; and calculating the sound source angle information of the location of the person speaking based on the vector value.

[0211] In some embodiments of this application, before obtaining the current shooting angle of the camera, the process includes: querying the current operating state of the camera; if the current operating state of the camera is in a rotating state, then waiting for the camera to finish rotating; if the current operating state of the camera is in a non-rotating state, then obtaining the current shooting angle of the camera.

[0212] In some embodiments of this application, determining the target rotation direction and target rotation angle of the camera based on the current shooting angle and sound source angle information of the camera includes: converting the sound source angle information into the coordinate angle of the camera; calculating the angle difference between the coordinate angle of the camera and the current shooting angle of the camera, and using the angle difference as the target rotation angle of the camera; and determining the target rotation direction of the camera based on the angle difference.

[0213] In some embodiments of this application, the step of converting the sound source angle information into the coordinate angle of the camera includes: obtaining the sound source angle range when the person is speaking and a preset angle range when the camera rotates; calculating the angle difference between the sound source angle range and the preset angle range, and taking half of the angle difference as the conversion angle; calculating the angle difference between the angle corresponding to the sound source angle information and the conversion angle, and taking the angle difference as the coordinate angle of the camera.

[0214] In some embodiments of this application, determining the target rotation direction of the camera based on the angle difference includes: if the angle difference is positive, then the target rotation direction of the camera is determined to be rightward; if the angle difference is negative, then the target rotation direction of the camera is determined to be leftward.

[0215] The second aspect:

[0216] In the embodiments of this application, such as Figure 15b As shown, the camera 232, acting as a detector 230, can be built into or connected to the external display device 200. After startup, the camera 232 can detect image data. The camera 232 can connect to the controller 250 via an interface component, thereby sending the detected image data to the controller 250 for processing. To detect images, the camera 232 may include a lens assembly and a pan-tilt assembly. The lens assembly can be an image acquisition element based on CCD (Charge Coupled Device) or CMOS (Complementary Metal Oxide Semiconductor) to generate image data of electrical signals based on the user image.

[0217] The lens assembly is mounted on the gimbal assembly, which can rotate the lens assembly to change its orientation. The gimbal assembly may include at least two rotating parts to respectively rotate the lens assembly left and right in a numerical direction and up and down in a horizontal direction. Each rotating part can be connected to a motor to drive its rotation automatically.

[0218] For example, such as Figure 17As shown, the gimbal assembly may include a first rotating axis in a vertical position and a second rotating axis in a horizontal position. The first rotating axis is located on the top of the display 275 and is rotatably connected to the top of the display 275. A fixing member is also provided on the first rotating axis, and the top of the fixing member is rotatably connected to the second rotating axis. The second rotating axis is connected to a lens assembly to drive the lens assembly to rotate. Motors and transmission components are respectively connected to the first and second rotating axes. The motors may be servo motors, stepper motors, etc., capable of automatically controlling the rotation angle. When a control command is received, the two motors can rotate respectively to drive the first and second rotating axes to rotate, thereby adjusting the orientation of the lens assembly.

[0219] By adjusting the orientation of the lens assembly, it can capture video of the user in different positions, thereby acquiring the user's image data. Clearly, different orientations correspond to image acquisition from different areas. When the user is positioned slightly to the left relative to the front of the monitor 275, the first pivot on the gimbal assembly can rotate the fixing component and the lens assembly to the left, ensuring the user's image is centered in the captured image. Conversely, when the user's body is positioned slightly lower, the second pivot on the gimbal assembly can rotate the lens assembly upwards, raising the shooting angle and centered the user's image in the image.

[0220] To track the position of a person, the controller 250 can identify the location of the user's image within the image by executing a person positioning and tracking method. If the user's position is unsuitable, the controller can rotate the camera 232 to obtain a suitable image. The identification of the user's position can be accomplished through image processing. For example, after activating the camera 232, the controller 250 can capture at least one image using the camera 232 as a calibration image. Feature analysis is then performed on the calibration image to identify the human image region. By determining the position of the human image region, the controller can determine whether the user's position is suitable.

[0221] However, in practical applications, the initial orientation of the camera 232 may be offset from the user's position in space. In some cases, the camera 232's shooting range may not cover the user's image, resulting in the camera 232 failing to capture the user's image or only capturing a small portion of it. This can lead to the inability to identify the image area during image processing and to effectively control the rotation of the camera 232 when the user's position is unsuitable; that is, it cannot effectively adjust for people not currently in the image.

[0222] Therefore, to ensure that the calibration image captured by camera 232 includes the human image area, the user's location can be determined using sound signals before acquiring the calibration image. After obtaining the location, camera 232 is rotated to face that location before acquiring the calibration image, thus making it easier for the acquired calibration image to include the human image area. For this purpose, display device 200 also includes a sound acquisition unit 231. The sound acquisition unit 231 can form an array using multiple microphones to simultaneously acquire the user's sound signals, thereby determining the user's location through the acquired sound signals. That is, as... Figure 18a , Figure 18b As shown, in some embodiments of this application, a method for locating and tracking a person with sound and image is provided, including the following steps:

[0223] Obtain the test audio signal input by the user.

[0224] In practical applications, the controller 250 can automatically run the audio-visual person positioning and tracking method after the camera 232 is activated, and acquire the test audio signal input by the user. The camera 232 can be activated manually or automatically. Manual activation means that the user selects the icon corresponding to the camera 232 in the operation interface using a control device such as a remote control 100. Automatic activation can occur automatically after the user performs certain interactive actions that require access to the camera 232. For example, if the user selects the "Mirror" application in the "My Applications" interface, since this application requires access to the camera 232, the camera 232 will also be activated simultaneously with the application's launch.

[0225] The camera 232 can be in the default initial posture after startup, for example, the default initial posture can be set to the lens assembly of the camera 232 facing forward; the posture after startup can also be the posture maintained when the camera 232 was last used. For example, if the camera 232 was adjusted to a 45-degree angled position when it was last used, then the posture of the camera 232 will also be a 45-degree angled position after startup this time.

[0226] After the camera 232 is activated, the controller 250 can acquire the test audio signal input by the user through the sound acquisition unit 231. Since the sound acquisition unit 231 includes a microphone array, microphones at different locations can acquire different audio signals for the same test audio.

[0227] In order to acquire audio signals through the microphone array, after the camera 232 is activated, text prompts can be automatically displayed on the display 275 and / or voice prompts can be played through audio output devices such as speakers to prompt the user to enter test audio, such as "Please enter test audio: Hi! Xiaoju".

[0228] It should be noted that the test audio can be various audio signals emitted by the user, including: voice produced by the user speaking, sounds produced by the user through body movements such as clapping, and sounds produced by the user through other handheld terminals. For example, when the user controls the display device 200 through a smart terminal such as a mobile phone, when the user needs to input a test audio signal, a control command can be sent to the smart terminal to control its sound output. This allows the smart terminal to automatically play a specific sound upon receiving the control command, so that the sound acquisition unit 231 can perform detection.

[0229] Therefore, in some embodiments, the controller 250 can acquire sound signals through a sound acquisition component after running the application, and extract voiceprint information from the sound signals. The voiceprint information is then compared with a preset test voiceprint. If the voiceprint information matches the preset test voiceprint, the sound signal is marked as a test audio signal; if the voiceprint information differs from the preset test voiceprint, the display 275 is controlled to display a prompt interface.

[0230] For example, when the test audio signal is set to the voice of the content "Hi! Xiaoju", after the microphone detects the sound signal, the voiceprint information in the sound signal can be extracted, and it can be determined whether the current voiceprint information is the same as the voiceprint information of "Hi! Xiaoju". After determining that the voiceprint information is the same, the subsequent steps are executed.

[0231] Clearly, this method of using smart terminals to generate sound can produce sounds with specific waveforms or loudness, giving the corresponding audio signals unique sound characteristics. This facilitates subsequent comparative analysis of the audio signals and mitigates the impact of other sounds in the environment on the analysis process.

[0232] The target location is determined based on the test audio signal.

[0233] After acquiring the test audio signal input by the user, the controller 250 can analyze the test audio signal to determine the target location of the user. Since the sound acquisition unit 231 includes multiple microphones forming a microphone array, the distance between different microphones and the sound source location varies relative to the sound source location, resulting in a certain propagation delay between the audio signals they acquire. The controller 250 can determine the approximate location of the user when they emitted the sound by analyzing the propagation delay between at least two microphones, combined with the distance between the two microphones and the speed of sound in the air.

[0234] By detecting time delays using multiple microphones, the location of the sound source can be determined, i.e., the target orientation can be identified. Since the purpose of detecting the target orientation is to orient the lens assembly of camera 232 towards that orientation, the target orientation can be represented solely by a relative angle. This allows controller 250 to directly determine the relative angle data after locating the target orientation and use it to calculate the angle that camera 232 needs to be adjusted. The relative angle can be the angle between the target position and the perpendicular line to the plane where camera 232 is located (i.e., the plane parallel to the display 275 screen), or the angle between the target position and the lens axis of camera 232.

[0235] For example, the external sound acquisition unit 231 connected to the display device 200 includes two microphones, respectively positioned on the two sides of the display 275, while the camera 232 is positioned at the center of the top edge of the display 275. When the user inputs a voice signal from either position, the microphones on both sides can detect the test audio signal, and then... Figure 19 From the positional relationship in the middle, we can know that:

[0236] Target orientation φ = arctan(L2 / D); where L2 is the horizontal distance between the user and camera 232, and D is the vertical distance between the user and camera 232.

[0237] According to the Pythagorean theorem, the following positional relationship can be determined: Monitor width H = L1 + L2 + L3; D 2 +(L1+L2) 2 =S1 2 L3 2 +D 2 =S2 2 Where S1 is the distance between the user's position and the left microphone, and S2 is the distance between the user's position and the right microphone, and S2 = vt; S1 = v(t + Δt), where v is the speed of sound in the air, t is the time it takes for the sound to reach the right microphone, and Δt is the time difference between the left and right microphones acquiring the test audio signal.

[0238] In the above equations, the display width H, the propagation speed v, and the acquisition time difference Δt are known. Therefore, L2 / D can be solved using the above positional relationships, and then the target orientation φ can be solved.

[0239] As can be seen, in this embodiment, the controller 250 can acquire test audio signals collected by at least two microphones and then extract the acquisition time difference of the test audio signals. Based on the acquisition time difference and the installation position data of the microphones and camera, the controller can calculate the target orientation. To obtain a more accurate target orientation, the positional relationship can be determined in both the horizontal and vertical directions, thereby calculating the horizontal and vertical deflection angles of the user's position relative to the camera's position. For example, the number of microphones can be increased or the microphones can be placed at different heights to determine the positional relationship in the vertical direction and calculate the vertical deflection angle.

[0240] It should be noted that the more microphones there are, the more accurately the user's location can be located, and the better the time delay between the audio signals received by different microphones can be detected. Therefore, in practical applications, the accuracy of target location detection can be improved by appropriately increasing the number of microphones. Furthermore, to increase the time delay and reduce detection error interference, the distance between the microphones can be increased to obtain even more accurate detection results.

[0241] The rotation angle is calculated based on the target orientation and the current posture of the camera 232.

[0242] After determining the direction from which the user emitted the sound (i.e., the target direction), the rotation angle of camera 232 can be calculated so that the camera's lens assembly can be oriented towards the target direction according to the rotation angle. For example, as... Figure 18a , Figure 18b As shown, the current camera 232 is in the default initial posture, and the relative angle between the target position and the vertical line of the screen is 30° to the left. Therefore, the calculated rotation angle φ is 30° (+30°) to the left.

[0243] Obviously, regardless of how the target's orientation is represented by a relative angle, the rotation angle can be calculated by the actual camera 232 based on its position and current posture. For example, if the current camera 232 is in a posture rotated 50° to the left, and the relative angle between the located target's orientation and the vertical line of the screen is 30° to the left, then the calculated rotation angle is 20° (-20°) to the right.

[0244] It should be noted that since the purpose of detecting the user's position through the test audio signal is to ensure that the calibration image captured by the camera 232 includes the user's corresponding facial image area, in most cases, controlling the rotation of the camera 232 in one direction will allow the captured calibration image to include the facial image area. However, in a few cases, such as when the current posture of the camera 232 is at its maximum vertical angle, rotating it in the horizontal direction will not allow the camera 232 to capture the facial image.

[0245] Therefore, in some embodiments, the target orientation in space (including the height direction) can also be determined by multiple microphones, and when calculating the rotation angle, the target orientation is decomposed into two angular components in the horizontal and vertical directions, thereby controlling the rotation angle of the camera 232 respectively.

[0246] A rotation command is generated based on the rotation angle, and the rotation command is sent to the camera 232.

[0247] After calculating the rotation angle, the controller 250 can encapsulate the rotation angle and generate a rotation command. This rotation command is then sent to the camera 232. Upon receiving the control command, the motor in the camera 232 can rotate, thereby driving the lens assembly to rotate via a pivot and adjusting the orientation of the lens assembly.

[0248] As can be seen from the above technical solution, the display device 200 can connect to an external camera 232 and a sound collector 231 through an interface component. After entering an application that requires human image tracking, the sound collector 231 collects test audio signals through multiple microphones and locates the target position of the user. This allows the camera 232 to rotate so that the lens component faces the user's position, thereby adjusting the shooting direction of the camera 232 to face the target position. This facilitates the acquisition of images containing the user's image, enabling adjustments even when there is no image area on the current screen, thus achieving subsequent human image tracking.

[0249] In order to achieve the tracking of people, after the camera 232 has completed its rotation, the controller 250 can continue to execute the audio-visual person positioning and tracking method. By acquiring images, the controller can identify the position of the person in the image and control the camera 232 to rotate to track the user's position when the position of the person changes, so that the person in the image captured by the camera 232 is always in the appropriate area.

[0250] Specifically, in some embodiments, after the camera 232 rotates to face the target according to the rotation command, the controller 250 can also acquire a calibration image through the camera 232 and detect a human figure pattern in the calibration image; then, by marking the human figure pattern and sending a tracking command to the camera 232 when the user moves, the controller can track the user's position. By tracking the user's position, the human figure pattern in the image captured by the camera 232 can always be in a suitable position, such as in the central area of ​​the image, thereby achieving a better display effect in the application interface when performing functions such as "mirror" and "motion follow".

[0251] To track the user's location, in some embodiments, the controller 250 can acquire calibration images via the camera 232 at a set frequency and detect the position of the human image within the calibration images. Depending on the image layout required by the application, different preset area ranges can be set according to the application type. When the human image is within the preset area, it means that the position of the human image in the currently acquired calibration image is appropriate, and the current shooting direction of the camera 232 can remain unchanged. When the human image is no longer within the preset area, it means that the user's current position has moved a large distance, the position of the human image in the acquired calibration image is inappropriate, and the shooting direction of the camera 232 needs to be adjusted.

[0252] Therefore, the controller 250 can generate a tracking command based on the position of the human image and send the tracking command to the camera 232 to control the camera 232 to adjust its shooting direction. Clearly, after the camera 232 receives the tracking command, the adjusted shooting direction should be able to keep the human image within a preset area. For example, the audio-visual person positioning and tracking method further includes the following steps:

[0253] Detect user location.

[0254] After rotating and adjusting the camera 232, the camera 232 can capture multiple frames of images in real time and send the captured images to the controller 250 of the display device 200. The controller 250 can perform image processing according to the launched application, such as controlling the display 275 to display the image; on the other hand, it can analyze the calibration image by calling the detection program to determine the user's location.

[0255] The user's location can be detected through image processing. Specifically, it involves capturing images in real-time from the camera 232 and detecting limb information. This limb information includes key points and the outline of the limb, and the detected key points and limb outline positions are used to determine their location within the image. Key points can refer to a series of points in a human image that represent human features, such as the eyes, ears, nose, neck, shoulders, elbows, wrists, waist, knees, and ankles.

[0256] Key points can be determined through image recognition. This involves analyzing the characteristic shapes in the image and matching them with a preset template to identify the corresponding image and its location. The location of each key point can then be represented by the number of pixels in the image relative to its boundary. Based on the resolution and viewing angle of the camera 232, a Cartesian coordinate system can be constructed with the top-left corner of the image as the origin and right and down as positive directions. Each pixel in the image can then be represented using this Cartesian coordinate system.

[0257] For example, such as Figure 20 As shown, the horizontal and vertical camera viewing angles are HFOV and VFOV, respectively. The viewing angle can be obtained from the camera's CameraInfo. The camera preview image supports 1080P, with a width of 1920 and a height of 1080 pixels. Therefore, the position of each pixel in the image can be (x, y), where the value range of x is (0, 1920) and the value range of y is (0, 1080).

[0258] To accurately represent the user's location, multiple keypoints are typically used. During a single detection process, the positions of all or part of these keypoints need to be extracted to determine the outline of the limb. For example, there might be 18 keypoints: 2 eye points, 2 ear points, 1 nose point, 1 neck point, 2 shoulder points, 2 elbow points, 2 wrist points, 2 waist (or hip) points, 2 knee points, and 2 ankle points. Clearly, these keypoints will require different recognition methods depending on the user's orientation. For instance, the position corresponding to the waist might be recognized as a waist point when the user is facing the monitor (275°), but as a hip point when the user is facing away from the monitor (275°).

[0259] Obviously, when the user's location or posture changes, the positions of some key points will also change. With this change, the relative position of the human body in the image captured by camera 232 will also change. For example, when the human body moves to the left, the human body's position in the image captured by camera 232 will be shifted to the left, making image analysis and real-time display difficult.

[0260] Therefore, after detecting the user's location, it is also necessary to compare the user's location with the preset area in the calibration image to determine whether the current user's location is within the preset area.

[0261] In some embodiments, the user's position can be represented by the center position of the limb bounding box, which can be calculated from the coordinates of the detected key points. For example, by obtaining the x-axis coordinates of the key points on the left and right sides of the limb bounding box, the center position of the limb bounding box can be calculated, i.e., the x-axis coordinate of the center position is x0 = (x1 + x2) / 2.

[0262] Since the camera 232 in this embodiment can include two rotations in the left-right direction and one in the up-down direction, after calculating the x-axis coordinate of the center position, the x-axis coordinate can be judged to determine whether the x-axis coordinate of the center position is located at the center of the entire image. For example, when calibrating the image as a 1080P image (1920, 1080), the horizontal coordinate of the center point of the calibrated image is 960.

[0263] After determining the center position of the person and the center point of the image, the user's position can be compared to determine whether it is within a preset judgment area. To avoid increasing the processing load due to frequent adjustments and to allow for some detection errors, an allowable coordinate range can be preset based on the actual application requirements and the horizontal viewing angle of the camera 232. When the center position of the person is within the allowable coordinate range, the current user position is determined to be within the preset area. For example, if the maximum allowable coordinate error is 300 pixels, then the allowable coordinate range is [660, 1260]. When the detected user center position coordinates are within this range, the user is determined to be within the preset judgment area, meaning the calculated center position coordinates are not significantly different from the 960 position; when the detected user center position coordinates are not within this range, the current user position is determined to be outside the preset area, meaning the calculated center position coordinates are significantly different from the 960 position.

[0264] After comparing the user's position with a preset area in the calibration image, it can be determined whether facial tracking is needed based on the comparison results. If the current user position is not within the preset area, the camera 232 is rotated to position the user's image in the center of the frame. If the current user position is within the preset area, there is no need to rotate the camera 232; simply maintaining the camera's orientation is sufficient.

[0265] When the current user position is not within the preset area, in order to control the camera 232 to rotate, the controller 250 can calculate the rotation angle based on the user position and generate a control command based on the rotation angle to control the camera 232 to rotate.

[0266] Specifically, after determining that the current user's position is not within the preset area, the controller 250 can first calculate the distance between the center position of the human image area and the center point of the image area; then, based on the calculated distance, combined with the maximum viewing angle of the lens assembly of the camera 232 and the image size, calculate the rotation angle; finally, send the calculated rotation angle to the camera 232 in the form of a control command, so that the motor in the camera 232 drives each rotating shaft to rotate, thereby adjusting the orientation of the lens assembly.

[0267] For example, such as Figure 21 , Figure 22 As shown, the preview resolution of camera 232 is 1920x1080, the horizontal width of the image is imgWidth = 1920; the horizontal center coordinate of the image is x = 960; the center coordinate of the human image area is (x0, y0), and the horizontal center coordinate is x0; the horizontal viewing angle is hfov; then the center distance between the human image area and the image area is hd = x – x0. The horizontal rotation angle of camera 232 can then be calculated using the following formula:

[0268]

[0269] The above formula can be used to calculate the angle that the camera 232 needs to be adjusted. The controller 250 then compares the coordinate values ​​of the center position of the portrait area with the center point of the image area to determine the orientation of the center position of the portrait area relative to the center point of the image area, thereby determining the rotation direction of the camera 232. That is, if the horizontal position of the center of the portrait area is larger than that of the image center, the camera 232 is rotated to the right; otherwise, the camera 232 is rotated to the left. In this embodiment, the camera 232 can adopt a rear camera mode, so that the image displayed on the screen and the image captured by the camera are mirror images, that is, the horizontal angle rotation is reversed left and right.

[0270] After determining the rotation angle and direction, the controller 250 can encapsulate the rotation angle and direction data, generate control commands, and send the control commands to the camera 232. Upon receiving the control commands, the motor in the camera 232 can rotate, thereby driving the lens assembly to rotate via the rotating shaft and adjusting the orientation of the lens assembly.

[0271] It should be noted that in the above embodiment, the judgment and adjustment are based on the horizontal coordinate. In practical applications, the lens assembly can also be adjusted by comparing the vertical difference between the center position of the portrait area and the center position of the image area. The specific adjustment method is the same as the horizontal adjustment method. That is, after determining that the current user position is not in the preset area, the controller 250 can first calculate the vertical distance between the center position of the portrait area and the center point of the image area; then, based on the calculated vertical distance, combined with the maximum vertical viewing angle of the lens assembly of the camera 232 and the image size, calculate the rotation angle; finally, send the calculated rotation angle to the camera 232 in the form of a control command, so that the motor in the camera 232 drives the second rotating shaft to rotate, thereby adjusting the orientation of the lens assembly.

[0272] However, in practical applications, due to the influence of user posture and the different requirements of different applications, using the center position as the method for determining the user's position in some application scenarios cannot achieve good display, detection, and tracking results. Therefore, in some embodiments, controlling the rotation of the camera 232 to place the user's imaging position in the central area of ​​the image can be performed according to the following steps.

[0273] The first recognition point is detected in the proofread image.

[0274] The first identification point is one or more key points identified to represent the position of a part of the user's limbs. For example, the first identification point can be two eye points (or two ear points) to represent the user's head position. By matching the regions corresponding to the eye patterns (or ear patterns) in the proofreading image, it is detected whether the current image contains the first identification points, i.e., whether it contains eye points (or ear points).

[0275] If the proofreading image does not contain a first identification point, a second identification point is detected in the proofreading image.

[0276] The second identification point is a key point that is spaced a certain distance from the first identification point and has a relative positional relationship. For example, the second identification point can be the chest point. Since the chest point is located below the eye point in normal use, and the distance between the chest point and the eye point is 20-30cm, the direction that needs to be adjusted can be determined by detecting the chest point.

[0277] If the second identification point is detected in the calibration image, the rotation direction is determined according to the positional relationship between the second identification point and the first identification point.

[0278] For example, if the first recognition point, i.e. the eye point, is not detected in the calibration image, but the second recognition point, i.e. the chest point, is detected, it is determined that the user's head image cannot be fully displayed in the current calibration image, and the camera 232 needs to be raised so that the human head enters the preset area of ​​the image.

[0279] Obviously, in practical applications, depending on the relative positions of the second and first recognition points, the determined rotation direction will differ when the first recognition point is not detected in the calibration image but the second recognition point is detected. For example, if the first recognition point is the waist and the second recognition point is the chest, and the waist point is not detected but the chest point is detected, it indicates that the captured image is too close to the upper half of the portrait. Therefore, the shooting angle can be lowered to bring the lower half of the portrait into the preset area of ​​the image.

[0280] The camera 232 is controlled to rotate according to the rotation direction and preset adjustment step size so that the human image is located in the preset area of ​​the image.

[0281] For example, if key points such as eyes / ears (first recognition points) are not detected, but key points such as shoulders (second recognition points) are detected, the camera 232 can be raised to adjust the position of the first recognition point by 100 pixels each time until the first recognition point is at the 1 / 7-1 / 5 position.

[0282] If the image being calibrated contains a first recognition point, then obtain the position of the first recognition point relative to the image region.

[0283] By recognizing the image within the proofreading image, if a first recognition point is identified, its location can be further extracted to determine its position relative to the entire image region. For example, ... Figure 23a As shown, after obtaining the calibrated image, if the eye point is identified, i.e., the first identification point is determined, the current coordinates P(x1, y1) of the eye point can be obtained. Then, the x-axis coordinate value and / or y-axis coordinate value in the current coordinates are compared with the overall width imgWidth and / or height imgHeight of the image to determine the position of the first identification point relative to the image area.

[0284] Specifically, the position of the first recognition point relative to the image region can be determined in both the horizontal and vertical directions. Specifically, in the horizontal direction, the position of the first recognition point relative to the image region is x1 / imgWidth; in the vertical direction, the position of the first recognition point relative to the image region is y1 / imgHeight.

[0285] After obtaining the position of the first recognition point relative to the image region, it is also possible to determine the interval in which the position of the first recognition point is located, and to determine different adjustment methods based on the different intervals in which it is located.

[0286] For example, such as Figure 23a As shown, when detecting the position of the first recognition point relative to the image area in the vertical direction, if the eye (or ear) is detected to be below 1 / 5 of the image height, the eye position is too low. The camera 232 needs to be lowered to raise the eye position to a suitable area. During the downward pressing of the camera 232, if the eye point is detected at 1 / 5 of the image height, the pressing stops, completing the adjustment of the camera 232. Figure 23b As shown. When the detected eye (or ear) position is between 1 / 7 and 1 / 5 of the image height, the current first recognition point position is determined to be appropriate. Therefore, the height of camera 232 does not need to be adjusted to prevent frequent camera movements caused by shaking.

[0287] The above embodiments, through a combination of image recognition and other methods, enable real-time control of the orientation of the camera 232, achieving the tracking of human targets. Clearly, in practical applications, sound source localization can also be used to track human targets. Therefore, in some embodiments of this application, the tracking of human targets can employ a combination of sound source localization and image recognition for more accurate target localization.

[0288] For example, when running fitness apps with large movements and fast motions, it is possible to identify specific time periods in advance through statistical methods that make it difficult to determine the user's location. During these periods, audio signals can be acquired to help determine the user's location. The results of image recognition and audio positioning at this time can be combined to improve the accuracy of tracking human targets.

[0289] Furthermore, in some usage scenarios, multiple human faces may be detected through image recognition, which will affect the tracking process of camera 232. Therefore, in some embodiments of this application, a locking procedure can be used to lock onto one human face among multiple human faces for tracking. For example, the human face closest to the center of the screen can be found within a certain area in the center of the screen as the optimal face information (the central 1 / 3 area of ​​the screen size, where it appears most frequently), and the person's information can be recorded and locked. If no face information is detected, it indicates that the sound information error is large, and the person closest to the screen is locked.

[0290] Once a person is locked onto a camera, the adjustment of camera 232 can be affected only by the position of the locked person. That is, movement of other people within the image captured by camera 232 will not adjust camera 232; camera 232 remains stationary. Only movement of the locked person, detected by image detection, will cause camera 232 to rotate to follow the locked person.

[0291] As can be seen from the above technical solution, the display device 200 can acquire a calibration image through the camera 232, detect a human image pattern in the calibration image, and mark the human image pattern. Furthermore, it can send a tracking command to the camera when the user moves to track the user's position, achieving the effect of the camera 232 following the user's movement. By tracking the user's position, the human image pattern in the image captured by the camera 232 can always be in a suitable position, facilitating display, retrieval, and analysis by the application.

[0292] In some embodiments, in the step of marking the portrait pattern, if the proof image includes multiple portrait patterns, the portrait pattern located in the central region of the proof image is searched; if the central region of the proof image contains a portrait pattern, the portrait pattern in the central region of the image is marked; if the central region of the proof image does not contain a portrait pattern, the portrait pattern with the largest area in the proof image is marked.

[0293] For example, the controller 250 can query the status of the camera 232 in real time. If the camera 232 stops rotating based on the test audio signal, the AI ​​image detection algorithm is activated. It searches for facial information of people located outside the center of the screen within a certain area, records the information, and locks onto the target. If no facial information is detected, it indicates a large error in the audio information, and the closest person to the screen is locked onto the target.

[0294] In some embodiments, before acquiring the test audio signal input by the user, image recognition can be performed on the image captured by the camera 232 to determine whether the camera 232 can capture an image with a human figure. If a human figure is identified from the captured image, target tracking can be performed directly through subsequent image processing without the need for sound source localization. That is, after activating the camera 232, an initial image for human figure recognition can be acquired first, and the human figure region can be identified in the initial image. The human figure region can be identified using the same method as in the above embodiments, i.e., by recognizing key points.

[0295] If the initial image contains a human figure area, the user's position is detected directly, and subsequent steps are performed to track the human figure target through image processing. If the initial image does not contain a human figure area, the user's input test audio signal is acquired, and subsequent steps are performed. The camera 232 is adjusted to face the user's position using sound source localization, and then the user's position is detected and subsequent steps are performed again.

[0296] To achieve more accurate human image location determination, in some embodiments, such as Figure 24a , Figure 24b As shown, after identifying multiple key points, a skeletal diagram can be created based on these key points, allowing for further determination of the person's location. The skeletal lines are defined by connecting multiple key points. The shape of the skeletal lines varies depending on the user's posture.

[0297] It should be noted that the drawn skeletal lines can also be used to dynamically adjust the camera's shooting position based on the movement and changes of the skeletal lines. For example, if the change in the skeletal line's movement is determined to be from a squatting position to a standing position, the camera's angle can be raised so that the person in the standing position is also within a suitable area of ​​the image, i.e., from... Figure 24a Transition to Figure 24b The effect shown is that, in determining the change in skeletal line movement from a standing to a squatting position, the viewing angle of camera 232 can be lowered so that the person in the squatting position can also be placed within a suitable area of ​​the image, i.e., from... Figure 24b Transition to Figure 24a The effect shown.

[0298] The above embodiment illustrates the tracking of a person by camera 232 using the example of the person's position being at the center of the image. It should be understood that, depending on actual needs, the person's position in the captured image may be located in areas other than the center region. For example, such as... Figure 25a As shown, for motion-following applications, the display device 200 can render a virtual coach image based on the video captured by the camera 232, so that the scene image viewed by the user through the display device 200 includes both the user's image and the virtual coach's image. In this case, for rendering to follow the scene, the image captured by the camera 232 needs to be positioned on one side of the image, while the other side is used to render the virtual coach image.

[0299] For example, such as Figure 25a , Figure 25b As shown, when the current portrait position is determined to be in the center area of ​​the image through image calibration, a rotation command is also needed to be sent to the camera 232 to rotate the camera 232 so that the portrait is located in the right area of ​​the image.

[0300] As can be seen from the above technical solutions, compared with methods that rely solely on image processing or sound source localization for person tracking, the audio-visual person localization and tracking method provided in this application can improve upon the shortcomings of low accuracy in sound source localization, which fails to effectively locate the specific position of a person, and poor spatial perception in image processing, which can only locate the area pointed at by the camera 232. The audio-visual person localization and tracking method comprehensively utilizes sound source localization and image analysis from the camera 232. Leveraging the strong spatial perception capability of sound source localization, it first confirms the approximate position of the person and drives the camera 232 to face the sound source. Simultaneously, utilizing the high accuracy of image analysis from the camera 232, it performs person detection in the captured image to determine the specific position, driving the camera to make fine adjustments, thereby achieving precise localization and ensuring that the person captured by the camera 232 is focused and displayed in the image.

[0301] Based on the above-described audio-visual person positioning and tracking method, in some embodiments, this application also provides a display device 200, including: a display 275, an interface component, and a controller 250.

[0302] The display 275 is configured to display a user interface, and the interface component is configured to connect the camera 232 and the sound collector 231. The camera 232 can rotate to capture an image, and is configured to capture an image. The sound collector 231 includes a microphone array consisting of multiple microphones and is configured to collect audio signals.

[0303] The controller 250 is configured to acquire a test audio signal input by the user and, in response to the test audio signal, locate the target orientation. The target orientation is calculated based on the time difference of the test audio signal acquired by the sound acquisition component, thereby sending a rotation command to the camera to adjust the camera's shooting direction to face the target orientation.

[0304] In the above embodiments, the camera 232 and the sound acquisition device 231 can be externally connected through the interface component and combined with the display device 200 to complete the above-mentioned sound and image person positioning and tracking method. In some embodiments, the camera 232 and the sound acquisition device 231 can also be directly built into the display device 200, that is, the display device 200 includes a display 275, a camera 232, a sound acquisition device 231, and a controller 250. The camera 232 and the sound acquisition device 231 can be directly connected to the controller 250, so that the test audio signal can be directly obtained through the sound acquisition device 231, and the camera 232 can be directly controlled to rotate, thereby completing the above-mentioned sound and image person positioning and tracking method.

[0305] In a specific implementation, this application also provides a computer storage medium, wherein the computer storage medium may store a program, and when the program is executed, it may include some or all of the steps of the various embodiments of the camera shooting angle adjustment method provided in this application. The storage medium may be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0306] Those skilled in the art will clearly understand that the techniques in the embodiments of this application can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of this application, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application or some parts of the embodiments.

[0307] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0308] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.

Claims

1. A display device, characterized in that, include: monitor; It is configured to present one or more images and one or more user interfaces, wherein the one or more images include images obtained from a broadcast system or network; An interface component is configured to connect a camera and a sound acquisition component, wherein the camera is rotatable and configured to capture images at different angles. To capture images; The sound acquisition component includes a microphone array consisting of multiple microphones, configured to acquire audio signals; The controller is configured to: initiate the acquisition of a test audio signal input by the user when the image captured by the camera does not contain a human figure; In response to the test audio signal, the target location is determined, the target location being acquired by the sound acquisition component. The time difference of the test audio signal is calculated; a rotation command is sent to the camera to adjust the camera's shooting direction to the target position; Acquire a calibration image until the image captured by the camera contains a human figure pattern, then stop acquiring the test audio input by the user again, and generate a tracking instruction based on the position of the human figure pattern in the calibration image, the position of the human figure pattern in the calibration image being determined based on a skeletal line graphic established by multiple key points identified in the calibration image; Human portrait patterns are detected in the calibrated image, and a preset region is determined; wherein, the preset region is set with a maximum allowable coordinate error based on the center position of the human portrait. If the portrait pattern is within the preset area, the camera's shooting direction is maintained; If the human image pattern is not within the preset area, the camera's shooting direction is adjusted in response to the tracking command.

2. The display device according to claim 1, characterized in that, The controller performs image acquisition and calibration until the image captured by the camera contains a human figure pattern, and is further configured as follows: The calibration image is acquired through the camera; Identify at least one key point in the proofread image, and create a skeletal line graphic based on the identified key point; The system determines and marks the position of the human figure based on the skeletal graph, and sends a tracking command to the camera when the user moves. The system also adjusts the camera's shooting direction based on the human figure's position to track the user's location.

3. The display device according to claim 2, characterized in that, The controller executes commands when the user moves. The camera sends a tracking command and adjusts its shooting direction according to the position of the person in order to track the user's position. It is further configured to track the user's position according to the following steps: Acquire calibration images via camera at a set frequency; Detect the position of the portrait pattern in the proofread image; If the human figure pattern is not within the preset area, a tracking instruction is generated based on the position of the human figure pattern. The tracking instruction includes the rotation direction and rotation angle. Send the tracking command to the camera.

4. The display device according to claim 2, characterized in that, The controller performs the action of marking the location of the human image. In this step, if the proofread image includes multiple portrait patterns, it is further configured as follows: Locate the human figure pattern located in the center region of the proofread image; If the human figure pattern is present in the central region of the calibrated image, mark the position of the human figure pattern corresponding to the human figure pattern in the central region of the image. If the central region of the proofreading image does not contain the portrait pattern, mark the position of the portrait pattern with the largest area in the proofreading image.

5. The display device according to claim 1, characterized in that, The controller executes the sending of rotation to the camera. The command to switch instructions is further configured as follows: The initial image is acquired through the camera; Identify human figures in the initial image; If the initial image contains the human figure pattern, a rotation command is sent to the camera; If the initial image does not contain the human figure pattern, then the test audio signal input by the user again for performing human positioning is acquired.

6. The display device according to claim 1, characterized in that, The controller, in response to the tracking command, adjusts the camera's shooting direction if the human image pattern is not within the preset area, and is further configured to: Obtain the skeletal line graphics from multiple frames of proof images; The user's movement state is identified based on the skeletal line pattern; The motion change pattern is calculated based on the motion state corresponding to multiple frames of calibrated images, and the shooting direction of the camera is dynamically adjusted according to the motion change pattern.

7. The display device according to claim 1, characterized in that, The controller, when the image captured by the camera does not contain a human figure, initiates the acquisition of a test audio signal input by the user, and is further configured to: The sound signal is acquired through the sound acquisition component; Extract voiceprint information from the sound signal; Compare the voiceprint information with the preset test voiceprint; If the voiceprint information is the same as the preset test voiceprint, the sound signal is marked as a test audio signal; If the voiceprint information is different from the preset test voiceprint, the display will show a prompt interface.

8. The display device according to claim 1, characterized in that, The controller executes the sending of rotation to the camera. The command to switch instructions is further configured as follows: Acquire a proofread image and detect the user's location in the proofread image; Compare the position of the person with the preset judgment area; If the image of the person is located within the preset judgment area, the display is controlled to show the image captured by the camera in real time. If the image location is outside the preset judgment area, calculate the coordinate difference between the image location and the center of the preset judgment area; A rotation command is generated based on the coordinate difference, and the rotation command is sent to the camera.

9. A display device, characterized in that, include: monitor; It is configured to present one or more images and one or more user interfaces, wherein the one or more images include images obtained from a broadcast system or network; A camera, which is rotatable and configured to capture images; The sound acquisition component, including a microphone array consisting of multiple microphones, is configured to acquire audio signals; The controller is configured as follows: When the image captured by the camera does not contain a human figure, the system initiates the acquisition of a test audio signal input by the user. In response to the test audio signal, the target location is determined, and the target location is calculated based on the time difference of the test audio signal acquired by the sound acquisition component. Send a rotation command to the camera to adjust the camera's shooting direction to the target position; The system acquires a calibration image until a human figure is captured in the image, at which point it stops acquiring test audio input by the user. It also generates tracking instructions based on the position of the human figure in the calibration image. The position of the pattern in the proofreading image is determined based on a skeletal line graphic established from multiple key points identified in the proofreading image; Human portrait patterns are detected in the calibrated image, and a preset region is determined; wherein, the preset region is set with a maximum allowable coordinate error based on the center position of the human portrait. If the image of the person is within the preset area, keep the camera's shooting direction unchanged; If the human image pattern is not within the preset area, the camera's shooting direction is adjusted in response to the tracking command. all.

10. A method for locating and tracking a person with sound and image, characterized in that, Applied to a display device, the display device including a display and a controller, the display being configured to present one or more images and one or more user interfaces, wherein the one or more images include images obtained from a broadcast system or network; The display device has a built-in camera and a sound acquisition component connected externally via an interface component. The camera can rotate to capture an angle. The sound and image person positioning and tracking method includes: When the image captured by the camera does not contain a human figure, the system initiates the acquisition of a test audio signal input by the user. In response to the test audio signal, the target location is determined, and the target location is calculated based on the time difference of the test audio signal acquired by the sound acquisition component. Send a rotation command to the camera to adjust the camera's shooting direction to the target position; Acquire a calibration image until the image captured by the camera contains a human figure pattern, then stop acquiring the test audio input by the user again, and generate a tracking instruction based on the position of the human figure pattern in the calibration image, the position of the human figure pattern in the calibration image being determined based on a skeletal line graphic established by multiple key points identified in the calibration image; Human portrait patterns are detected in the calibrated image, and a preset region is determined; wherein, the preset region is set with a maximum allowable coordinate error based on the center position of the human portrait. If the image of the person is within the preset area, keep the camera's shooting direction unchanged; If the human image pattern is not within the preset area, the camera's shooting direction is adjusted in response to the tracking command.