A camera interaction control method, device, equipment and medium

CN122802778APending Publication Date: 2026-09-22SHENZHEN XIAOPAI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610964501.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]本发明实施例提供一种摄像头交互控制方法、装置、设备及介质,以解决现有智能摄像头交互方式单一、用户无法主动实时与智能摄像头进行高效直观交互的问题

Benefits of technology

[0008]上述摄像头交互控制方法、装置、计算机设备及存储介质的技术方案中,摄像头交互控制方法适用于投影仪,包括步骤:获取用户语音指令,并对语音指令进行语义解析,获得语义信息;接收预设的摄像头采集并传输的视频画面;将语义信息和视频画面进行融合分析,生成摄像头的设备控制指令,并将设备控制指令发送至摄像头;接收摄像头执行设备控制指令后返回的指令执行结果,并将指令执行结果通过投影画面和/或语音反馈给目标用户。该方法通过投影仪将用户语音指令与实时视频画面进行深度融合分析,实现了对摄像头的智能化、直观化控制,有效提升了用户与摄像头之间的交互效率和体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802778A_ABST
    Figure CN122802778A_ABST
Patent Text Reader

Abstract

The application discloses a camera interactive control method, device, equipment and medium, the camera interactive control method is suitable for a projector, including the following steps: obtaining a user voice instruction, and performing semantic analysis on the voice instruction to obtain semantic information; receiving a preset video picture collected and transmitted by a camera; fusing and analyzing the semantic information and the video picture to generate a device control instruction of the camera, and sending the device control instruction to the camera; receiving an instruction execution result returned after the camera executes the device control instruction, and feeding back the instruction execution result to a target user through a projection picture and / or voice. The method fuses and analyzes the user voice instruction and the real-time video picture in depth through the projector, realizes intelligent and intuitive control of the camera, and effectively improves the interactive efficiency and experience between the user and the camera.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent security technology, and in particular to a camera interactive control method, device, equipment, and medium. Background Technology

[0002] Smart cameras are among the most widely used smart devices in modern homes and offices, with increasingly diverse applications. However, traditional smart cameras primarily rely on passive triggering mechanisms (such as motion detection and sound detection) to activate recording and alarms. Their interaction methods are limited, preventing users from actively and in real-time interacting with the camera. Furthermore, when detecting and locating targets using a smart camera, users mainly rely on manually operating a mobile app or computer client to view the footage and control the pan / tilt mechanism. This lack of efficient and intuitive interaction between the user and the camera results in a poor user experience. Summary of the Invention

[0003] This invention provides a camera interaction control method, device, equipment, and medium to solve the problem that existing smart cameras have limited interaction methods and users cannot actively and intuitively interact with smart cameras in real time.

[0004] In a first aspect, this application provides a camera interactive control method applicable to a projector, comprising the steps of: acquiring a user's voice command and performing semantic parsing on the voice command to obtain semantic information; receiving a preset video image captured and transmitted by a camera; fusing and analyzing the semantic information and the video image to generate a device control command for the camera, and sending the device control command to the camera; receiving the command execution result returned by the camera after executing the device control command, and feeding back the command execution result to the target user through the projected image and / or voice.

[0005] Secondly, this application provides a camera interactive control device suitable for projectors, comprising: a semantic parsing module for acquiring user voice commands and performing semantic parsing on the voice commands to obtain semantic information; an image receiving module for receiving video images captured and transmitted by a preset camera; a fusion analysis module for fusing and analyzing the semantic information and video images to generate device control commands for the camera and sending the device control commands to the camera; and a result feedback module for receiving the command execution result returned by the camera after executing the device control command and feeding back the command execution result to the target user through the projected image and / or voice.

[0006] Thirdly, this application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described camera interaction control method.

[0007] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described camera interaction control method.

[0008] In the aforementioned technical solutions for camera interaction control methods, devices, computer equipment, and storage media, the camera interaction control method is applicable to projectors and includes the following steps: acquiring user voice commands and performing semantic parsing on the voice commands to obtain semantic information; receiving video images captured and transmitted by a preset camera; fusing and analyzing the semantic information and video images to generate device control commands for the camera, and sending the device control commands to the camera; receiving the command execution result returned by the camera after executing the device control commands, and feeding back the command execution result to the target user through the projected image and / or voice. This method deeply fuses and analyzes user voice commands with real-time video images through a projector, achieving intelligent and intuitive control of the camera, effectively improving the interaction efficiency and experience between the user and the camera. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a flowchart of a camera interaction control method according to an embodiment of the present invention; Figure 2 This is a detailed flowchart of step S3 in the camera interaction control method of one embodiment of the present invention; Figure 3 This is another specific flowchart of step S3 in the camera interaction control method in one embodiment of the present invention; Figure 4 This is a detailed flowchart of a remote security event playback between a user terminal and a camera in a camera interaction control method according to an embodiment of the present invention; Figure 5 This is a schematic diagram of a camera interaction control device according to an embodiment of the present invention; Figure 6 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0012] First, it should be noted that in the following embodiments of the camera interaction control method provided in this application, the projector and the camera are wirelessly connected using StarFlash technology. The projector acts as the local master control terminal, and the camera uses at least one IPC (Internet Protocol Camera). In this end-to-end collaborative architecture with the IPC camera as the execution terminal, the two are directly connected via StarFlash, achieving real-time end-to-end interactive control without the need for a router or the internet. Specifically, StarFlash technology, as a novel short-range wireless communication technology, offers low latency, allowing voice command transmission latency to be controlled at the microsecond level. Its high bandwidth ensures real-time transmission of high-definition video images, and its anti-interference capability guarantees stable connections in complex electromagnetic environments. The introduction of this technical architecture provides a reliable data link foundation for subsequent in-depth fusion analysis.

[0013] The projector, serving as the main control unit, integrates a first-speech communication module, a microphone array, a local speech recognition module, a local semantic parsing module, a multimodal fusion decision module, and a projection display module. The first-speech communication module establishes a low-latency, high-bandwidth wireless direct link with the camera; the microphone array collects user voice commands, supporting multi-angle sound source localization and noise reduction; the local speech recognition module and the local semantic parsing module work together to perform offline speech-to-semantics conversion; and the multimodal fusion decision module is responsible for spatiotemporally aligning and semantically associating the voice command parsing results with the video feed from the camera, generating deep fusion analysis results, and driving the projection display module for visualization.

[0014] Meanwhile, the camera, as the execution end, integrates a second satellite communication module, an image sensor, a video encoding module, a target detection module, a PTZ control module, an electronic zoom module, and a local storage module. The second satellite communication module establishes a direct wireless link with the projector's first satellite communication module; the image sensor captures high-definition video images; the video encoding module efficiently compresses and encodes the raw video data captured by the image sensor to meet the bandwidth requirements of the satellite link; the target detection module identifies and locates targets of interest in the video image in real time, supporting the detection and tracking of various targets such as people, vehicles, and objects; the PTZ control module drives the camera to perform precise horizontal and vertical rotation based on control commands issued by the fusion analysis module; the electronic zoom module performs optical zoom magnification on specific areas of the image based on user commands or automatic detection results, ensuring clear visibility of details; and the local storage module caches keyframe images and target detection results, supporting offline playback and abnormal event retrieval.

[0015] In one embodiment, such as Figure 1 As shown, a camera interaction control method is provided, including the following steps: Step S1: Obtain the user's voice command and perform semantic parsing on the voice command to obtain semantic information.

[0016] It should be noted that this step is a local offline semantic parsing process, requiring no cloud computing power or uploading of voice data, thus achieving complete local closed-loop processing of voice interaction data. User voice commands are natural language, spoken commands input by the user towards the projector, distinct from traditional fixed-format commands, supporting everyday, informal, and ambiguous statements. Semantic parsing refers to the projector's process of converting voice signals into text information and performing deep semantic understanding, utilizing its locally deployed local speech recognition and semantic parsing modules. Semantic information is the deep understanding result obtained after semantic parsing the voice commands, including user intent and command parameters. User intent refers to the user's control needs regarding the IPC camera, specifically including core interactive intents such as zooming in on the screen area, detecting a specific target, gimbal rotation, image tracking, and gimbal full-area scanning search. Command parameters are the limiting conditions matching the corresponding user intent, specifically including controllable parameters such as spatial orientation words, target object words, gimbal scanning angle, and dwell time.

[0017] In this embodiment, the user's voice commands are collected through the microphone array built into the projector, and the local speech recognition module ASR (Automatic Speech Recognition) and the local semantic parsing module LLM (Large Language Model) are used to perform deep semantic understanding on the voice commands to obtain semantic information containing the user's intent and command parameters.

[0018] For example, if a user inputs the voice command "zoom in on the bottom right corner of the screen," after this step's parsing, the user's intent is identified as zooming in on a region, and the corresponding command parameter is the orientation parameter of the bottom right corner of the screen. If a user inputs the voice command "find the key," after this step's parsing, the user's intent is identified as object detection, and the corresponding command parameter is the target to be detected, "key."

[0019] This application utilizes local offline acquisition of user voice commands and precise parsing to extract semantic information containing user intent and command parameters, abandoning the traditional interaction mode that relies on manual operation via a mobile app and cloud-based voice parsing for camera interaction. On one hand, it eliminates dependence on networks and cloud servers, enabling normal voice interaction control even in offline environments, significantly lowering the interaction threshold. On the other hand, it achieves local processing of voice data, effectively avoiding the privacy risks associated with cloud transmission of voice data, while providing accurate and effective semantic input for subsequent multimodal fusion decision-making based on voice and vision, and precise control of camera actions.

[0020] Step S2: Receive the video footage captured and transmitted by the preset camera.

[0021] It should be noted that in this step, the camera acts as the execution end, and its image sensor collects the on-site monitoring images in real time. The video encoding module compresses the original video images using H.265 encoding to reduce the bandwidth usage of video transmission. The high-definition video data stream after encoding is transmitted to the projector in real time through the Star Flash wireless link established between the second Star Flash communication module and the first Star Flash communication module.

[0022] The SparkLink wireless link utilizes the SparkLink Basic (SLB) mode, which boasts a peak data transmission rate of ≥12Mbps, stably supporting real-time transmission of 1080p@30fps high-definition video. The corresponding transmission latency is ≤0.25ms, ensuring end-to-end transmission delay is within 50ms. This guarantees real-time, smooth, and stutter-free video reception by the projector, eliminating issues such as image delay, frame drops, and blurriness. It supports up to 4096 connected devices and simultaneous access from multiple cameras. Indoor coverage ranges from 50-100 meters, meeting the needs of most households. Furthermore, SparkLink technology supports point-to-point direct communication between devices, eliminating the need for routers, LANs, and internet access. This enables full-featured offline interaction even without a network connection, completely freeing traditional camera interaction devices from network dependence. In addition, the StarScan physical layer and channel adopt a dual-layer encryption mechanism, which can realize local encrypted transmission of video data, prevent video data from being eavesdropped on or tampered with, ensure the security of monitoring screen transmission, and provide a solid communication foundation for local end-side closed-loop interaction.

[0023] Furthermore, after the projector receives the real-time video image transmitted by the IPC camera through the second flash communication module via the first flash communication module, it caches and stores it locally. This provides real-time and complete visual data support for subsequent multimodal fusion analysis of semantic information and video images, as well as the generation of control strategies. It enables accurate matching of voice semantic commands with real-time images, laying a data foundation for subsequent precise control of the camera.

[0024] Step S3: Perform fusion analysis on semantic information and video footage to generate device control commands for the camera, and send the device control commands to the camera.

[0025] It should be noted that the core of this step is the multimodal fusion decision-making and instruction generation process. Among them, fusion analysis refers to the projector performing cross-modal association matching between the text semantic information obtained from offline parsing and the real-time acquired visual video images, thereby resolving natural language ambiguities and completing semantic scene adaptation. Device control instructions refer to the standardized control signals generated by the projector based on the fusion analysis results, which can be recognized and executed by the IPC camera hardware, and are used to drive hardware actions such as camera zoom, gimbal rotation, target detection, and full-area scanning.

[0026] In this embodiment, the projector binds user intent and command parameters from semantic information with the screen coordinates and content features of the real-time video image. It performs fusion analysis according to different scenarios, generates corresponding device control commands, and sends them to the IPC camera via the StarSpark SLB direct connection link, achieving accurate conversion of voice commands into camera hardware actions. The overall system is divided into two core interactive scenarios: area zoom and target detection, each corresponding to different sub-step execution logic.

[0027] In this application, by using multimodal fusion analysis of semantics and video images, the traditional control mode of mechanical execution of single voice commands without scene verification is changed. By relying on real monitoring images to correct the ambiguity of spoken commands, the intelligent interactive effect of "understanding commands, understanding scenes, and precise control" is achieved. This effectively solves the problems of inaccurate control and misoperation caused by the ambiguity of natural language commands, and greatly improves the intelligence and accuracy of camera interaction.

[0028] In some embodiments, when the user intent includes area zooming and the instruction parameters include spatial orientation words, the device control instructions include camera zoom control instructions and / or gimbal rotation control instructions.

[0029] Specifically, such as Figure 2 As shown, step S3 includes the following sub-steps: Step S301: Map the spatial location words to the screen coordinate system of the video frame to determine the location coordinates of the target area that needs to be magnified in the video frame.

[0030] It should be noted that spatial directional words refer to directional descriptive words used in user voice commands to limit the position of the screen, including colloquial directional parameters such as upper left corner, lower right corner, left side, and below; the screen coordinate system refers to a standardized pixel coordinate system established based on the real-time video screen received by the projector, used to accurately locate local areas of the screen.

[0031] Among them, the projector presets the mapping rules between the image orientation and pixel coordinates, accurately converts the colloquial spatial orientation words obtained from semantic parsing into the pixel coordinate range corresponding to the video image, locks the target area that needs to be magnified, and completes the accurate matching of voice orientation semantics to visual image coordinates.

[0032] In this application, the standardization and quantitative analysis of non-standard spoken directional commands are achieved through the mapping and transformation between directional words and the image coordinate system. This solves the problem of fuzzy natural language directional descriptions and the inability to directly drive hardware zoom, providing quantitative data support for precise local magnification and regional focusing.

[0033] Step S302: Generate zoom control commands and / or gimbal rotation control commands based on the area location coordinates. The zoom control commands and / or gimbal rotation control commands are used to magnify the target area for display.

[0034] It should be noted that the zoom control command is the control signal that controls the camera's electronic zoom module to adjust the focal length and magnify a specified area of ​​the image. The gimbal rotation control command is the control signal that controls the camera's gimbal to fine-tune the angle and center the target area for display. The two can be used individually or in combination.

[0035] The projector determines the position of the target area in the current screen based on the location coordinates of the locked target area. If the target area is within the field of view, only zoom control commands are generated. If the target area deviates from the center of the screen or partially goes out of the screen, gimbal rotation control commands and zoom control commands are generated simultaneously and then sent to the camera for execution via the StarFlash link.

[0036] In this application, by generating combined control commands based on the position and status of the screen, the focus and magnification of a specified area can be accurately completed without the need for the user to manually adjust the camera. This achieves automated screen focusing driven by natural language, greatly simplifying the operation process for detailed viewing of surveillance footage.

[0037] In some embodiments, when the user intent includes target detection and the instruction parameters include the term "target object", the device control instruction includes a target selection instruction from the camera or a gimbal scanning control instruction.

[0038] Specifically, such as Figure 3 As shown, step S3 includes the following sub-steps: Step S311: Based on the target object word, call the preset target detection model in the camera to perform target detection in the video frame and obtain the target detection result.

[0039] It should be noted that the target object word refers to the name of the object to be detected in the user's voice command, such as key, mobile phone, person, etc. The object detection model is a lightweight YOLO detection model deployed on the IPC camera side, which supports local real-time object recognition without the need for cloud computing power.

[0040] The projector sends the target object words obtained from semantic parsing to the IPC camera, calls the lightweight target detection model on the camera side to perform frame-by-frame recognition and detection of the current real-time video screen, and matches whether there is a target object in the screen that corresponds to the target object words.

[0041] Furthermore, the target detection model can adopt lightweight network structures such as YOLOv5s or YOLOv8n, which can control the number of model parameters to the MB level while ensuring detection accuracy, meet the MB-level computing power requirements, adapt to the limited computing power resources of IPC cameras, and realize real-time inference on the edge.

[0042] In this application, a lightweight detection model on the edge is used to achieve local real-time detection without uploading the image data to the cloud. While ensuring the real-time performance of target detection, local closed-loop processing of video image data is achieved, eliminating the risk of privacy leakage and balancing detection efficiency and data security.

[0043] Step S312: When the target detection result is that a matching target is detected in the video frame, a target selection instruction is generated and sent to the camera. The target selection instruction is used to control the camera to select and highlight the matching target in the video frame.

[0044] It should be noted that the target selection command is a dedicated control signal that drives the camera to mark the target position and lock the target image area. It can realize the visual highlighting of the target, making it easy for users to intuitively identify the detection results.

[0045] Specifically, after the camera detects a matching target, it sends the target pixel coordinates back to the projector. The projector generates a target selection command based on the target coordinates and sends it out to control the camera to mark the matching target with a border and highlight it in the video.

[0046] In this application, the target detection results are visualized by automatically selecting and labeling the target after successful detection, eliminating the need for users to manually identify the target in the image, and greatly improving the intuitiveness and interactive experience of intelligent detection.

[0047] In some embodiments, if multiple matching targets are detected from the video footage, further target filtering and priority ranking are required by incorporating visual features. Specifically, based on the confidence score output by the target detection model, multiple matching targets can be sorted from high to low confidence, with the target with the highest confidence score being selected first and designated as the primary tracking object. Simultaneously, matching targets with confidence scores below a preset threshold are marked as candidate tracking objects, and their motion trajectories are continuously monitored for backup, ensuring continuous tracking even in scenarios involving target movement or occlusion.

[0048] Step S313: When the target detection result is that no matching target is detected from the video image, generate a PTZ scanning control strategy.

[0049] It should be noted that the PTZ scanning control strategy is a full-domain detection rule system adapted to scenarios where targets are missing. The PTZ scanning control strategy includes the camera's preset scanning path, multiple preset scanning areas, and preset dwell time for each preset scanning area. The preset scanning path includes key parameters such as the scanning start position (e.g., horizontal angle -90 degrees, vertical angle -45 degrees), horizontal step angle (e.g., 30 degrees), vertical step angle (e.g., 20 degrees), scanning direction (e.g., horizontal S-shaped reciprocating scan, first left then right, then right to left), and scanning coverage (e.g., horizontal angle range 0 to 360 degrees, vertical angle range -45 to 90 degrees), ensuring that the camera can systematically traverse the monitored area according to predetermined logic. The preset scanning area is the camera's field-of-view division unit, used to divide the panoramic monitoring into several independent detection sub-regions. The preset dwell time is a fixed duration for single-region image acquisition and target detection, ensuring sufficient image acquisition and feature analysis are completed within a single region, avoiding missed detections due to excessively rapid scanning.

[0050] In this embodiment, the preset scanning path is preferably an S-shaped reciprocating scanning path, which can cover the entire field of view of the camera without blind spots; the preset scanning area is the equally divided field of view interval by the rotation of the camera gimbal; the preset dwell time is a fixed duration for single-area image acquisition and target detection to ensure detection accuracy.

[0051] Specifically, after receiving the target detection results of the camera that do not match the target, the projector retrieves the preset global scanning parameters and generates a standardized PTZ scanning control strategy that includes a preset scanning path with S-shaped reciprocating scanning, a preset scanning area with multi-partition scanning, and a preset dwell time of 1-2 seconds for a single preset scanning area, thus avoiding the tedious operation of manual scanning.

[0052] In this application, a full-domain scanning strategy is automatically generated for scenes where there is no target in the current frame. This overcomes the technical deficiency of traditional cameras, which can only detect the current frame and cannot detect targets after they leave the frame. This achieves full-domain coverage of target detection and greatly improves the completeness and automation of intelligent detection.

[0053] Step S314: Based on the PTZ scanning control strategy, generate PTZ scanning control instructions. The PTZ scanning control instructions are used to control the camera to perform scanning and searching according to the PTZ scanning control strategy.

[0054] It should be noted that the PTZ scanning control command is an integrated control signaling that encapsulates complete scanning strategy parameters, which can directly drive the camera PTZ to operate automatically and perform zone detection according to preset rules.

[0055] In this embodiment, the projector encapsulates the scanning path, area parameters, and dwell time in the PTZ scanning control strategy to generate standardized PTZ scanning control commands, which are then sent to the IPC camera at high speed via the StarFlash link to initiate the full-area automatic scanning and detection process.

[0056] In this application, the entire detection process can be started and stopped automatically through the issuance of integrated scanning commands, without the need for manual intervention, thereby further reducing the operational threshold of intelligent monitoring and improving the intelligence level of the equipment.

[0057] Furthermore, the method also includes the steps of: controlling the camera to move along a preset scanning path based on the gimbal scanning control command, sequentially covering each preset scanning area with the camera's field of view, and staying in each preset scanning area for a preset dwell time, and calling the target detection model to perform target detection in the corresponding preset scanning area.

[0058] Specifically, based on the PTZ scanning control command, the camera is controlled to move at a constant speed along a preset S-shaped reciprocating scanning path, so that the camera's field of view can be fully covered in each preset equally divided scanning area in sequence. The camera stays in each preset scanning area for a preset dwell time of 1-2 seconds. During the dwell time, the lightweight target detection model on the end side is called to perform accurate target detection on the current frame. The full field of view detection is completed section by section, and the entire monitoring scene is covered without blind spots.

[0059] It should be noted that the full-area scanning detection uses a closed-loop detection logic of zone-by-zone, timed pause, and frame-by-frame detection, which is different from the traditional gimbal constant-speed cruise. This can ensure the detection accuracy of each field of view area and avoid missed detections and false detections.

[0060] Specifically, after receiving the pan-tilt-zoom (PTZ) scanning control command, the camera drives the PTU to rotate in an S-shaped reciprocating path, sequentially traversing all preset scanning areas. After reaching each scanning area, it pauses for a preset dwell time and simultaneously starts the edge-side target detection model to complete the target detection of the current area. This process is repeated until the entire area is scanned.

[0061] In this application, the scanning logic of zoned stopping and zone-by-zone detection avoids the problems of blurred images and missed targets caused by dynamic cruising, ensuring the accuracy of full-area scanning detection and realizing intelligent object search without blind spots.

[0062] When a matching target is detected in any preset scanning area, the camera's gimbal scanning motion is stopped, and a target selection command and a zoom control command are generated. The target selection command and zoom control command are used to select and zoom out the matching target.

[0063] It should be noted that the target selection command and zoom control command are the linkage control logic after the target is detected, which can quickly lock the target, focus the image, and complete the detection loop.

[0064] Specifically, once a matching target is detected in any area during the scanning process, the camera immediately stops the pan-tilt scanning motion, locks the current field of view, and transmits the target coordinates back to the projector; the projector simultaneously generates a target selection command and a zoom control command, which are then sent to the camera for execution.

[0065] In this application, the linkage logic of immediately stopping scanning upon target detection and focusing and magnifying enables rapid positioning and accurate display, thereby improving the response efficiency and visualization effect of full-domain detection.

[0066] Furthermore, if no matching target is detected in any of the preset scanning areas, a voice prompt indicating that the target does not exist is generated.

[0067] It should be noted that the voice prompt indicating that the target does not exist is a fallback feedback mechanism in the detection loop, used to inform the user of the overall detection results and avoid the user from waiting continuously.

[0068] After the camera completes the detection of the entire scanning area and fails to detect a matching target, it sends the status of no detection results back to the projector. The projector then generates a corresponding voice prompt and plays the voice prompt indicating that the target does not exist through its local voice broadcast module.

[0069] In this application, full-domain detection with voice feedback is used to improve the intelligent detection interaction loop, allowing users to quickly know the detection results and improving the completeness of the interaction and user experience.

[0070] Step S4: Receive the command execution result returned after the camera executes the device control command, and feed back the command execution result to the target user through the projected image and / or voice.

[0071] It should be noted that the command execution result refers to the status data returned by the camera after completing hardware actions such as zooming, gimbal rotation, target selection, and global scanning. The command execution result includes target detection results and command execution status. Target detection results include target coordinates, target image features, and target type recognition results; command execution status includes gimbal rotation status, zoom level, scanning progress, and selection status. User feedback refers to the closed-loop interaction mechanism by which the projector presents the interactive results to the user through visual displays and voice announcements based on the execution results.

[0072] After completing the hardware operations corresponding to various device control commands, the camera transmits the complete command execution results back to the projector in real time via the StarFlash bidirectional link. The projector receives the execution results and combines them with the real-time monitoring screen. The projection display module integrates the target detection results to generate a visual feedback screen for projection display, and generates voice prompts to indicate the command execution status. User feedback is then simultaneously provided through the projection screen and the local voice broadcast module.

[0073] Furthermore, when the target detection result indicates that a matching target has been detected in the video frame, the bounding box of the matching target is displayed on the projected screen, and the command execution status is fed back to the target user through voice broadcast.

[0074] For example, for the user voice command "zoom in on the bottom right corner of the screen," the camera receives the zoom control command, performs the zoom operation, and transmits the zoomed image back to the projector. The projector simultaneously projects the updated display and announces, "The bottom right corner of the screen has been zoomed in." For the user voice command "find the key in the image," the camera performs target detection and bounding box selection, transmitting the detected target results, including target coordinates and image features, back to the projector. The projector overlays the target detection results and bounding box marks onto the projected image and announces, "The key has been found." If no key is detected, the camera sends "No key detected" back to the projector, which generates the corresponding pan-tilt-zoom (PTZ) scanning control command to control the camera to perform a full-area scan. If the full-area scan still does not detect the key, the camera announces "No key found" to the user, completing the detection loop. If a target is detected in a preset scanning area, the pan-tilt scanning operation stops, the camera performs target detection and bounding box selection, and the detected target detection results, including the target coordinates and image features, are sent back to the projector. The projector overlays the target detection results and bounding box marks onto the projected screen, and announces "The key has been found" to the target user via voice, while simultaneously bounding boxing and marking the target location on the projected screen.

[0075] In this application, a closed-loop interactive system of "voice command input - intelligent decision control - equipment execution - result feedback" is constructed by transmitting the results of command execution and providing feedback in multiple forms. This allows users to intuitively and quickly grasp the results of camera control, solving the problems of no feedback and unintuitive results in traditional monitoring control, and greatly improving the integrity and user experience of intelligent interaction.

[0076] It should be noted that the video recordings captured by the camera adopt a hybrid storage architecture, mainly including three storage levels: local full-video recording storage, event recording storage, and thumbnail storage. Local full-video recording storage saves all video data continuously captured by the camera, supports playback along the timeline, and is stored on an IPC TF card. It can be directly accessed and played back via a projector. Event recording storage saves key video clips that trigger security events, stored on an IPC TF card, and supports local viewing by event type or remote retrieval via a temporary channel. Thumbnail storage saves compressed images of key frames corresponding to security events, stored on user terminals (such as mobile phones), and can be pushed via low-power networks to achieve low-power, low-bandwidth real-time alarm notifications.

[0077] Furthermore, the method also includes a remote security event playback process between the user terminal and the camera. This process is independent of the local StarFlash interaction process. Under the premise of strictly ensuring the local video privacy closed loop, it enables low-risk and controllable remote security event viewing and on-demand recording playback. The entire process follows the privacy and security logic of "local storage, no cloud access, abbreviated alarms, on-demand retrieval, time-limited encryption, and destruction after viewing".

[0078] Specifically, such as Figure 4 As shown, the camera interaction control method further includes the following steps: Step S51: Receive the security event thumbnail pushed by the camera through the target user's user terminal.

[0079] It should be noted that security events refer to pre-set security anomalies such as personnel intrusion, abnormal image movement, and area anomalies identified by the camera in real-time monitoring of scene changes; security event thumbnails are extremely low-resolution preview images generated by the camera by compressing key frames of security events, with a data size of only tens of KB, used only for event reminders, lacking complete image details, and unable to reconstruct the monitored scene; the user terminal is the user's bound mobile device, which only establishes a low-power communication link with the IPC camera, and the complete video data is not relayed through a cloud server throughout the process.

[0080] During local monitoring, the IPC camera performs real-time security detection. Once a security anomaly is identified and triggered, the complete high-definition video clip corresponding to the event is immediately stored locally on the device's built-in TF card. Uploading, backing up, or synchronizing any original video data to the cloud server is strictly prohibited, achieving 100% local closed-loop storage of high-definition video data. Simultaneously, the camera extracts only keyframes of the event to generate lightweight security thumbnails. These thumbnails, along with brief text alarm information, are pushed point-to-point to the bound target user terminal via a low-power network. The user terminal receives and displays the security alarm thumbnail and event notification in real time.

[0081] In this application, by differentiating between "thumbnail alerts and local retention of original videos" through a layered push mechanism, the inherent mode of uploading all video data to the cloud in traditional security systems is changed. This not only allows users to remotely perceive abnormal events on-site in the first instance, but also avoids the risk of privacy leakage caused by the outflow of complete monitoring video and cloud retention from the source, thus balancing remote security early warning capabilities with local data privacy and security.

[0082] Step S52: In response to the target user's video playback request, establish an encrypted channel between the user terminal and the camera, and retrieve the original video data corresponding to the security event thumbnail through the encrypted channel.

[0083] It should be noted that the video playback request is a targeted video retrieval command initiated manually by the user after the user terminal displays the thumbnail alarm. The system does not proactively push or automatically upload the original video by default. The encrypted channel is a one-time dedicated encrypted transmission link temporarily created by the camera, corresponding only to the current user session. It has an independently negotiated session key and automatically expires upon channel expiration or after transmission ends; it cannot be reused. The original video data is stored locally on the camera's TF card and is associated with the alarm thumbnail. Figure 1 A complete high-definition video clip of the corresponding event.

[0084] In this system, after the user terminal receives and displays a security thumbnail alarm, it only sends a playback command to the IPC camera when the user actively clicks to play back or initiates a video retrieval request. The camera listens for the request signal from the remote user terminal in real time. After successful verification, it immediately establishes an independent point-to-point encrypted transmission channel between the user terminal and the camera, without any forwarding through a cloud server. Subsequently, the two parties dynamically negotiate and generate a unique session key. The camera reads the corresponding original security video from the local TF card and transmits it encrypted through the encrypted channel using the session key, allowing the user terminal to decrypt and complete local playback. When the video data transmission is complete, the user ends the playback operation, or any condition triggers the 5-minute validity period of the channel, the system automatically and completely destroys the current session key and permanently closes the temporary encrypted channel. After the channel is closed, the video data cannot be retrieved, viewed, or recovered through this channel again, achieving on-demand, one-time transmission of video data that becomes invalid after viewing.

[0085] In this application, a hybrid architecture combining thumbnail push and user-initiated retrieval is employed to meet remote viewing needs while ensuring privacy. Simultaneously, a high-security remote access system, distinct from traditional security systems, is constructed through a full-link privacy playback mechanism involving "user-initiated authorization request + point-to-point encrypted transmission from the terminal camera + time-limited validity + key destruction + channel failure." This ensures that high-definition recording is never uploaded to the cloud, never permanently stored, and never leaked, with only single-time controllable transmission after user authorization. This effectively addresses the technical pain points of traditional remote playback, such as long-term network video transmission, cloud-based recording storage, and the vulnerability of data to theft and misuse. Combined with the front-end projector's offline multimodal local interaction and network-free controllable operation architecture, a complete, secure, and intelligent camera interaction system of "local fully offline intelligent interaction + remotely controllable privacy playback" is formed, significantly improving the overall system's practicality and privacy security.

[0086] In summary, this application provides a camera interaction control method applicable to projectors, comprising the following steps: acquiring user voice commands and performing semantic parsing on the voice commands to obtain semantic information; receiving video images captured and transmitted by a preset camera; fusing and analyzing the semantic information and video images to generate device control commands for the camera, and sending the device control commands to the camera; receiving the command execution result returned by the camera after executing the device control commands, and feeding back the command execution result to the target user. This method deeply fuses and analyzes user voice commands with real-time video images through a projector, achieving intelligent and intuitive control of the camera, effectively improving the interaction efficiency and experience between the user and the camera. Users can complete complex camera control through natural language and view it directly on a large projection screen without needing to take out their mobile phones or open an app, thus breaking down the interaction barriers of traditional security equipment that rely on mobile apps, and realizing a leap from "passive monitoring" to "proactive intelligent service." Simultaneously, the projector, as the main control terminal, fully utilizes its computing power advantages and large-screen display characteristics to localize the complex visual perception and logical judgment processes, solving the ambiguity and vagueness of natural language commands, and ensuring accurate recognition and execution of control intentions. Furthermore, the projector is normally used for watching movies and entertainment, and can be switched to monitoring display for security purposes, improving equipment utilization and achieving multi-purpose functionality. This significantly enhances the equipment's value and space utilization in home settings. Meanwhile, the camera's video stream is processed entirely locally in a closed loop, without uploading to any cloud server. StarFlash physical layer encryption and channel encryption provide dual protection, fundamentally eliminating the risk of video data leakage during transmission and storage. This effectively safeguards user privacy and data sovereignty, and to some extent reduces investment costs and maintenance burdens.

[0087] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0088] In one embodiment, a camera interaction control device is provided, suitable for a projector, and this camera interaction control device corresponds one-to-one with the camera interaction control method described in the above embodiments. For example... Figure 5 As shown, the camera interaction control device includes a semantic parsing module 101, an image receiving module 102, a fusion analysis module 103, and a result feedback module 104. Detailed descriptions of each functional module are as follows: The semantic parsing module 101 is used to acquire user voice commands and perform semantic parsing on the voice commands to obtain semantic information.

[0089] The video receiving module 102 is used to receive video images captured and transmitted by a preset camera.

[0090] The fusion analysis module 103 is used to fuse and analyze semantic information and video images to generate device control commands for the camera and send the device control commands to the camera.

[0091] The result feedback module 104 is used to receive the instruction execution result returned after the camera executes the device control instruction, and to feed back the instruction execution result to the target user through the projected image and / or voice.

[0092] The semantic parsing module 101 includes the microphone array, local speech recognition module, and local semantic parsing module of the aforementioned projector. The image receiving module 102 corresponds to the first satellite communication module of the aforementioned projector. The fusion analysis module 103 corresponds to the multimodal fusion decision module of the aforementioned projector. The result feedback module 104 corresponds to the aforementioned projection display module.

[0093] Specific limitations regarding the camera interaction control device can be found in the limitations of the camera interaction control method described above, and will not be repeated here. Each module in the aforementioned camera interaction control device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0094] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a camera interactive control method.

[0095] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the camera interaction control method described in the above embodiment, for example... Figure 1 As shown in S1-S4, or Figures 2 to 4 As shown, to avoid repetition, it will not be described again here. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in this embodiment of the camera interaction control device, for example... Figure 5The functions of the semantic parsing module 101, the image receiving module 102, the fusion analysis module 103, and the result feedback module 104 shown are not described again here to avoid repetition.

[0096] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the camera interaction control method described in the above embodiment, for example... Figure 1 As shown in S1-S4, or Figures 2 to 4 As shown, to avoid repetition, it will not be described again here. Alternatively, when the computer program is executed by the processor, it implements the functions of each module / unit in this embodiment of the camera interaction control device, for example... Figure 5 The functions of the semantic parsing module 101, the image receiving module 102, the fusion analysis module 103, and the result feedback module 104 shown are not described again here to avoid repetition. The computer-readable storage medium can be non-volatile or volatile.

[0097] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0098] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0099] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.

[0100] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A camera interaction control method, applicable to projectors, characterized in that, Including the following steps: Obtain user voice commands and perform semantic parsing on the voice commands to obtain semantic information; Receive video footage captured and transmitted by a preset camera; The semantic information and the video footage are fused and analyzed to generate device control commands for the camera, and the device control commands are sent to the camera. The system receives the instruction execution result returned by the camera after executing the device control command, and feeds back the instruction execution result to the target user through a projected image and / or voice.

2. The camera interaction control method according to claim 1, characterized in that, The semantic information includes user intent and command parameters; when the user intent includes region zoom and the command parameters contain spatial orientation words, the device control commands include zoom control commands for the camera and / or gimbal rotation control commands. The step of fusing and analyzing the semantic information and the video footage to generate the device control commands for the camera includes: Map the spatial directional words to the screen coordinate system of the video frame to determine the location coordinates of the target area that needs to be magnified in the video frame. The zoom control command and / or gimbal rotation control command are generated based on the location coordinates of the area, and the zoom control command and / or gimbal rotation control command are used to magnify the display of the target area.

3. The camera interaction control method according to claim 1, characterized in that, The semantic information includes user intent and instruction parameters; when the user intent includes target detection and the instruction parameters contain target object words, the device control instruction includes the target bounding box instruction of the camera; the step of fusing and analyzing the semantic information and the video frame to generate the device control instruction of the camera includes: Based on the target object word, the preset target detection model in the camera is invoked to perform target detection in the video frame and obtain the target detection result; When the target detection result indicates that a matching target has been detected in the video frame, a target selection instruction is generated and sent to the camera. The target selection instruction is used to control the camera to select and highlight the matching target in the video frame.

4. The camera interaction control method according to claim 3, characterized in that, The device control commands also include pan-tilt scanning control commands. The step of fusing and analyzing the semantic information and the video footage to generate the camera's device control commands further includes: When the target detection result is that the matching target is not detected in the video frame, a gimbal scanning control strategy is generated; the gimbal scanning control strategy includes the camera's preset scanning path, multiple preset scanning areas, and preset dwell time for each preset scanning area; Based on the gimbal scanning control strategy, the gimbal scanning control command is generated, which is used to control the camera to perform scanning and searching according to the gimbal scanning control strategy.

5. The camera interaction control method according to claim 4, characterized in that, After generating the gimbal scanning control command based on the gimbal scanning control strategy, the method further includes: Based on the gimbal scanning control command, the camera is controlled to move along the preset scanning path, sequentially covering each preset scanning area with the camera's field of view, and staying in each preset scanning area for a preset time, while the target detection model is invoked to perform target detection in the corresponding preset scanning area; When the matching target is detected in any of the preset scanning areas, the pan-tilt scanning motion of the camera is stopped, and the target selection instruction and zoom control instruction are generated. The target selection instruction and zoom control instruction are used to select and zoom out the matching target.

6. The camera interaction control method according to claim 3, characterized in that, The instruction execution result includes the target detection result and the instruction execution status of the target selection instruction; The step of feeding back the execution result of the instruction to the target user via a projected image and / or voice includes: When the target detection result indicates that a matching target has been detected in the video frame, the frame of the matching target is displayed on the projected screen, and the execution status of the instruction is fed back to the target user via voice broadcast.

7. The camera interaction control method according to claim 1, characterized in that, The camera interaction control method further includes the following steps: The target user receives a thumbnail of a security event pushed by the camera through their user terminal. In response to the target user's video playback request, an encrypted channel is established between the user terminal and the camera, and the original video data corresponding to the security event thumbnail is retrieved through the encrypted channel.

8. A camera interactive control device, suitable for projectors, characterized in that, include: The semantic parsing module is used to acquire user voice commands and perform semantic parsing on the voice commands to obtain semantic information; The video receiving module is used to receive video images captured and transmitted by a preset camera; The fusion analysis module is used to fuse and analyze the semantic information and the video footage, generate device control commands for the camera, and send the device control commands to the camera; The result feedback module is used to receive the instruction execution result returned by the camera after executing the device control command, and to feed back the instruction execution result to the target user through a projected image and / or voice.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the camera interaction control method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the camera interaction control method as described in any one of claims 1 to 7.