Equipment control method and device, electronic equipment, storage medium and computer program product
By generating target images and enabling interactive operations within those images, the control method for PTZ cameras solves the problems of cumbersome operation and inaccurate positioning, achieving intuitive device control and efficient user interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN LUMIUNITED TECH CO LTD
- Filing Date
- 2025-10-21
- Publication Date
- 2026-04-28
AI Technical Summary
The control method of PTZ cameras relies on manual operation or angle parameter input, which makes operation cumbersome and positioning inaccurate, making it difficult to meet the high-precision, low-latency positioning requirements in complex scenarios.
By displaying a target image and a composite image generated from multiple original images, users can interact with the target image, obtain the mapping relationship between the interaction position and the marked position, deduce the target motion information, and control the device to move to the target position.
It improves the intuitiveness of device control and the efficiency of user interaction, avoids cumbersome angle input operations, and significantly improves the operating experience and positioning accuracy of PTZ camera equipment.
Smart Images

Figure CN121940573A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet of Things (IoT) technology, and more specifically, to a device control method, apparatus, electronic device, storage medium, and computer program product. Background Technology
[0002] With the rapid development of IoT technology, PTZ cameras, due to their horizontal and vertical rotation control capabilities and ability to cover a wider field of view, have become an important part of the video surveillance field. To improve monitoring efficiency, multiple preset positions are usually set in PTZ cameras to enable quick switching and viewing of specific areas.
[0003] However, most current methods for setting preset positions for PTZ cameras rely on manual operation or position jumps achieved by inputting angle parameters. The operation process is complex and lacks direct correlation with image content, resulting in low positioning accuracy.
[0004] As can be seen from the above, the control method of PTZ cameras still has problems such as cumbersome operation and inaccurate positioning. Summary of the Invention
[0005] This application provides a device control method, apparatus, electronic device, and storage medium, which can solve the problems of cumbersome operation and inaccurate positioning in the control methods of PTZ cameras in related technologies. The technical solutions are as follows: According to one aspect of this application, a device control method includes: displaying a target image; the target image being obtained based on multiple frames of original images captured by a target device; in response to a trigger operation on the target image, acquiring an interaction position corresponding to the trigger operation and a marker position corresponding to each of the original images, so as to obtain target motion information corresponding to the interaction position based on the mapping relationship between the interaction position and the marker positions and motion information of the target device; controlling the target device to move to a target position based on the target motion information, and displaying the image captured by the target device at the target position.
[0006] According to one aspect of this application, a device control apparatus includes: an image display module for displaying a target image; the target image is obtained based on multiple frames of original images captured by a target device; an information acquisition module for responding to a trigger operation on the target image, acquiring an interaction position corresponding to the trigger operation and a marker position corresponding to each of the original images, so as to obtain target motion information corresponding to the interaction position according to the mapping relationship between the interaction position and the marker positions and the motion information of the target device; and a device control module for controlling the target device to move to a target position based on the target motion information, and displaying the image captured by the target device at the target position.
[0007] In an exemplary embodiment, the device control device is further configured to acquire multiple frames of original images captured by the target device under different motion information; and to perform image synthesis processing on the multiple frames of original images to obtain the target image.
[0008] In an exemplary embodiment, the device control apparatus is further configured to assign corresponding marker positions to each of the original images based on motion information corresponding to the multiple frames of original images; and generate a target image carrying each of the marker positions according to the marker positions corresponding to each of the original images.
[0009] In an exemplary embodiment, the information acquisition module is further configured to acquire at least one marker position carried by the target image; and determine the marker position in the target image corresponding to the interaction position.
[0010] In an exemplary embodiment, the information acquisition module is further configured to identify the marker positions corresponding to each of the original images in the target image, and obtain the position of each marker position in the target image.
[0011] In an exemplary embodiment, the information acquisition module is further configured to find at least one adjacent marker position in the target image based on the interaction position; and based on the at least one adjacent marker position and the interaction position, according to the distance between each of the original images in the image synthesis direction, obtain the marker position whose distance meets a set condition, so as to determine the marker position in the target image corresponding to the interaction position.
[0012] In an exemplary embodiment, the information acquisition module is further configured to, when there are corresponding first and second marker positions at the interaction position, search for first motion information corresponding to the first marker position and second motion information corresponding to the second marker position according to the interaction position and the mapping relationship between each marker position and the motion information of the target device; and determine the target motion information based on the interaction position, the first motion information and the second motion information.
[0013] In an exemplary embodiment, the information acquisition module is further configured to calculate target motion information corresponding to the interaction position based on the relative position ratio between the interaction position and the first and second marker positions, as well as the first and second position information.
[0014] In an exemplary embodiment, the information acquisition module is further configured to, when there is a corresponding marker position at the interaction position, search for motion information corresponding to the marker position based on the interaction position and the mapping relationship between each marker position and the motion information of the target device; and use the motion information corresponding to the marker position as the target motion information corresponding to the interaction position.
[0015] According to one aspect of this application, an electronic device includes at least one processor and at least one memory, wherein the memory stores a computer program that, when executed by the processor, implements the device control method as described above.
[0016] According to one aspect of this application, a storage medium having a computer program stored thereon, which, when executed by one or more processors, implements the device control method as described above.
[0017] According to one aspect of this application, a computer program product includes a computer program that, when executed by one or more processors, implements the device control method as described above.
[0018] The beneficial effects of the technical solution provided in this application are: In the above technical solution, a target image is synthesized from multiple frames of original images captured by the target device. This allows users to directly perform click, touch, and other trigger operations on the target image. Based on the mapping relationship between the interaction position and each marked position and the motion information of the target device, the target motion information corresponding to the interaction position is derived, thereby controlling the target device to move to the target position. Through the above process of mapping interaction position, marked position, and motion information, the user's intuitive interactive behavior can be transformed into target motion information that can be used to control the target device, avoiding cumbersome operations such as relying on angle input and preset paths. This solution improves the intuitiveness of device control, significantly improves the operating experience of the target device and the efficiency of user interaction, and solves the problems of cumbersome operation and inaccurate positioning of PTZ camera devices during user control. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram based on the implementation environment involved in this application; Figure 2This is a hardware structure diagram of a terminal according to an exemplary embodiment; Figure 3 This is a flowchart illustrating a device control method according to an exemplary embodiment; Figure 4 yes Figure 3 A schematic diagram illustrating a specific implementation of acquiring raw images according to a corresponding embodiment; Figure 5 yes Figure 4 The three original images acquired by the target device involved in the corresponding embodiment; Figure 6 yes Figure 4 A schematic diagram illustrating a specific implementation of generating a target image according to a corresponding embodiment; Figures 7 to 8 yes Figure 3 A schematic diagram illustrating a specific implementation of an original image with marked positions and its synthesized target image, as described in the corresponding embodiment; Figure 9 yes Figure 3 A schematic diagram illustrating a specific implementation of an interactive position and a marker position in a corresponding embodiment; Figure 10 yes Figure 3 A schematic diagram illustrating another specific implementation of the interaction position and marker position involved in the corresponding embodiment; Figure 11 This is a schematic diagram illustrating the specific implementation of a device control method in an application scenario; Figure 12 This is a schematic diagram illustrating the processing of the original image in one embodiment; Figure 13 This is a schematic diagram of the processing of the original image in another embodiment; Figure 14 A structural block diagram of a device control apparatus according to an exemplary embodiment is shown; Figure 15 This is a structural block diagram of an electronic device according to an exemplary embodiment. Detailed Implementation
[0021] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0022] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this disclosure means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0023] As mentioned earlier, most current methods for setting preset positions for PTZ cameras rely on manual operation or position jumps achieved by inputting angle parameters. The operation process is complex and lacks direct correlation with image content, resulting in low positioning accuracy.
[0024] Specifically, the preset position of a PTZ camera is usually set in the following two ways: (1) the user manually controls the camera to rotate to the desired observation position and manually saves the current PTZ angle as a preset position; (2) the user directly inputs the corresponding horizontal angle (pan) and vertical angle (tilt) values in the system interface, thereby instructing the device to jump to the specified rotation position. Although this method achieves remote positioning and viewing angle adjustment at the functional level, it has significant usage barriers and operational inconvenience.
[0025] First, the operation process requires a high level of spatial awareness from the user, who needs to accurately judge the angle range of each target in the monitored area or confirm it through repeated trial and error, which is inefficient. Second, there is no direct semantic relationship between the preset angle information and the image content in the camera screen, making it difficult for users to quickly locate the target based solely on the screen content. Third, if there are multiple similar areas in the scene (such as multiple corridors, multiple devices, etc.), it is often difficult to accurately distinguish them by simply relying on the angle setting, which can easily lead to positioning errors or misselection of the location.
[0026] Therefore, the current gimbal control method based on manual settings or angle parameter input is difficult to meet the high-precision, low-latency positioning requirements in complex scenarios. Especially in visual monitoring systems that require fast and intuitive responses to user operation intentions, its positioning efficiency and user-friendliness still need to be further improved.
[0027] As can be seen from the above, the control methods for PTZ cameras still suffer from drawbacks such as cumbersome operation and inaccurate positioning.
[0028] Therefore, the device control method provided in this application can effectively simplify the operation of the pan-tilt camera and improve positioning accuracy. Accordingly, the device control method is applicable to device control devices, which can be deployed on electronic devices. The electronic devices can also refer to portable mobile electronic devices, such as smartphones, tablets, etc.
[0029] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0030] Figure 1 This is a schematic diagram of an implementation environment involved in a device control method. The implementation environment includes at least a user terminal 110, a smart device 130, a server 170, and network equipment. Figure 1 In this context, network devices include gateway 150 and router 190, but this is not intended to be a specific limitation.
[0031] The user terminal 110, which can also be considered as a user terminal or terminal, can deploy (or install) the client associated with the smart device 130. This user terminal 110 can be an electronic device such as a smartphone, tablet, laptop, desktop computer, smart control panel, or other device with display and control functions, and is not limited here.
[0032] The client, associated with the smart device 130, is essentially where the user registers an account and configures the smart device 130. For example, the configuration includes adding a device identifier to the smart device 130, so that when the client runs on the user terminal 110, it can provide the user with functions such as device display and device control of the smart device 130. This client can be in the form of an application or a web page. Correspondingly, the interface for displaying the device on the client can be in the form of a program window or a web page, and there is no limitation here.
[0033] Smart device 130 is deployed in gateway 150 and communicates with gateway 150 through its own configured communication module, thereby being controlled by gateway 150. It should be understood that smart device 130 generally refers to one of multiple smart devices 130. This application embodiment only uses smart device 130 as an example; that is, this application embodiment does not limit the number or type of smart devices deployed in gateway 150. In one application scenario, smart device 130 is deployed in gateway 150 by accessing it through a local area network. The process of smart device 130 accessing gateway 150 through a local area network includes: gateway 150 first establishes a local area network, and smart device 130 joins the local area network established by gateway 150 by connecting to it. This local area network includes, but is not limited to, ZIGBEE or Bluetooth. Among them, the smart device 130 can be a smart printer, smart fax machine, smart camera, smart air conditioner, smart door lock, smart light, or a human body sensor, door and window sensor, temperature and humidity sensor, water immersion sensor, natural gas alarm, smoke alarm, wall switch, wall socket, wireless switch, wireless wall sticker switch, cube controller, curtain motor, millimeter wave radar, etc., equipped with a communication module.
[0034] The interaction between user terminal 110 and smart device 130 can be achieved through a local area network (LAN) or a wide area network (WAN). In one application scenario, user terminal 110 establishes a wired or wireless communication connection with gateway 150 via router 190, such as Wi-Fi, allowing user terminal 110 and gateway 150 to be deployed on the same LAN, thus enabling user terminal 110 to interact with smart device 130 via the LAN path. In another application scenario, user terminal 110 establishes a wired or wireless communication connection with gateway 150 via server 170, such as 2G, 3G, 4G, 5G, or Wi-Fi, allowing user terminal 110 and gateway 150 to be deployed on the same WAN, thus enabling user terminal 110 to interact with smart device 130 via the WAN path.
[0035] The server-side 170 can also be considered as the cloud, cloud platform, platform side, server side, etc. This server-side 170 can be a single server, a server cluster consisting of multiple servers, or a cloud computing center consisting of multiple servers, in order to better provide backend services to a massive number of user terminals 110. For example, backend services include device control services.
[0036] In one application scenario, the target device can be a smart device 130, which displays a target image on a user terminal 110. The target image is obtained based on multiple frames of original images captured by the target device. In response to a trigger operation on the target image, the interaction position corresponding to the trigger operation and the marker position corresponding to each original image are obtained. Based on the mapping relationship between the interaction position and each marker position and the motion information of the target device, the target motion information corresponding to the interaction position is obtained. Based on the target motion information, the target device is controlled to move to the target position and the image captured by the target device at the target position is displayed.
[0037] Please see Figure 2 , Figure 2 This is a hardware structure diagram of a terminal according to an exemplary embodiment. The terminal is suitable for... Figure 1 The user terminal 110 in the implementation environment is shown.
[0038] It should be noted that this terminal is merely an example adapted to this application and should not be construed as providing any limitation on the scope of use of this application. Furthermore, this terminal should not be interpreted as requiring or depending on any specific feature. Figure 2 One or more components of the exemplary terminal 100 shown.
[0039] like Figure 2 As shown, terminal 100 includes memory 101, memory controller 103, and one or more ( Figure 2 (Only one is shown) Processor 105, peripheral interface 107, radio frequency module 109, positioning module 111, camera module 113, audio module 115, touch screen 117, and button module 119. These components communicate with each other through one or more communication buses / signal lines 121.
[0040] The memory 101 can be used to store computer programs, such as the computer program corresponding to the device control method and apparatus in the exemplary embodiment of this application. The processor 105 reads the computer program stored in the memory 101 to perform various functions and data processing, that is, to complete the device control method.
[0041] The memory 101, as a carrier for resource storage, can be random access memory, such as high-speed random access memory, non-volatile memory, such as one or more magnetic storage devices, flash memory, or other solid-state memory. The storage method can be temporary storage or permanent storage.
[0042] The peripheral interface 107 may include at least one wired or wireless network interface, at least one serial-to-parallel conversion interface, at least one input / output interface, and at least one USB interface, etc., for coupling various external input / output devices to the memory 101 and the processor 105 to realize communication with various external input / output devices.
[0043] The radio frequency module 109 is used to transmit and receive electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, thereby enabling communication with other devices through a communication network. The communication network includes cellular telephone networks, wireless local area networks (WLANs), or metropolitan area networks (MANs), and these networks can use various communication standards, protocols, and technologies.
[0044] The positioning module 111 is used to obtain the current geographical location of the terminal 100. Examples of positioning modules 111 include, but are not limited to, Global Positioning System (GPS), positioning technologies based on wireless local area networks or mobile communication networks.
[0045] The camera module 113 is part of the camera and is used to capture pictures or videos. The captured pictures or videos can be stored in the memory 101 or transmitted to the host computer via the radio frequency module 109.
[0046] The audio module 115 provides an audio interface to the user, which may include one or more microphone jacks, one or more speaker jacks, and one or more headphone jacks. Audio data is exchanged with other devices through the audio interface. Audio data can be stored in the memory 101 and can also be transmitted via the radio frequency module 109.
[0047] The touchscreen 117 provides an input / output interface between the terminal 100 and the user. Specifically, the user can perform input operations through the touchscreen 117, such as clicking, touching, and swiping gestures, so that the terminal 100 can respond to the input operations. The terminal 100 then displays the output content, which is any form or combination of text, images, or videos, to the user through the touchscreen 117.
[0048] The button module 119 includes at least one button, providing an interface for users to input information into the terminal 100. Users can press different buttons to enable the terminal 100 to perform different functions. For example, the volume adjustment button allows users to adjust the volume of the sound played by the terminal 100.
[0049] Understandable. Figure 2 The structure shown is for illustrative purposes only; the terminal 100 may also include components that are more advanced than those shown. Figure 2 The more or fewer components shown, or having the same Figure 2 The different components are shown. Figure 2 The components shown can be implemented using hardware, software, or a combination thereof.
[0050] Please see Figure 3 This application provides a device control method, which is applicable to electronic devices. For example, the electronic device may be... Figure 1The user terminal 110 shown in the implementation environment can have the following hardware structure: Figure 2 As shown.
[0051] In the following method embodiments, for ease of description, the execution subject of each step of the method is an electronic device, but this does not constitute a specific limitation.
[0052] like Figure 3 As shown, the method may include the following steps: Step 310: Display the target image.
[0053] The target image is obtained based on multiple frames of original images captured by the target device. That is, the target image can be a synthetic image generated from the original images and is interactive with the user. For example, it can be a panoramic image formed by stitching together multiple frames of original images, projection mapping, image enhancement, and other processing methods. The target image can support user interaction operations such as clicking and dragging to control the movement behavior of the target device.
[0054] The target device can refer to an electronic device with image acquisition capabilities and motion control support, such as a smart camera device or panoramic shooting device equipped with a camera module and gimbal mechanism. This target device can be used to acquire image data encompassing a specific space and perform viewing angle adjustments or position changes based on user interaction commands.
[0055] The raw image can refer to one or more frames of image data captured by the target device, meaning the original image frame that has not undergone fusion processing. Each raw image frame corresponds to an image captured by the target device from a specific viewpoint or position.
[0056] The target image can be displayed on a user terminal that is connected to the target device, such as a smartphone, tablet, or touchscreen device. Users can view and interact with the target image through the graphical user interface of this terminal, such as clicking, dragging, and zooming, to control the target device.
[0057] Step 330: In response to a trigger operation on the target image, obtain the interaction position corresponding to the trigger operation and the marker position corresponding to each original image, so as to obtain the target motion information corresponding to the interaction position based on the mapping relationship between the interaction position and each marker position and the motion information of the target device.
[0058] The triggering operation can refer to the user's interactive behavior at any location on the displayed target image, such as clicking, touching, or pressing. This type of triggering operation can be detected, and its corresponding coordinates can be extracted as the interaction location for subsequent calculation of target device control information.
[0059] When a user wants to control the direction or view of the target device, they can perform a trigger operation at a certain position in the target image to obtain the interaction position. The interaction position can be regarded as the coordinates corresponding to the user's trigger operation in the target image.
[0060] For example, in a user interface displaying a target image, if a user taps the screen slightly to the left of the center of the target image, the trigger operation can be recognized, and the position of the tap point in the image coordinate system can be obtained as the interaction position.
[0061] It should be noted that the marker position corresponding to each original image can be the synthetic position of the original image in the target image determined during the process of generating the target image based on multiple frames of original images acquired by the target device.
[0062] Specifically, when stitching multiple original images into a complete target image according to the image acquisition order or changes in device viewing angle, the starting position of each original image in the target image is recorded, such as the center point of the image. This position serves as the marker position corresponding to that original image frame.
[0063] For example, if the target device captures three original images, namely the left view, middle view, and right view, then when synthesizing the target image (such as a panoramic image), these three original images may be stitched sequentially into the left, middle, and right regions of the target image. In this case, recording the stitching start position or center position of the three images in the target image will form three sets of marked positions.
[0064] Then, a mapping relationship can be established with the motion information corresponding to the original image acquisition of each frame (such as the spatial coordinate information, viewpoint parameters, lens orientation parameters, etc. of the target device), thereby obtaining the mapping relationship between each marker position and the motion information of the target device.
[0065] It should be noted that in some cases, the interaction location may completely coincide with a marker location in the target image. In such cases, the motion information corresponding to the marker location can be directly used as the target device motion information corresponding to the interaction location, without further calculation.
[0066] However, in many real-world scenarios, due to overprocessing during the original image stitching process and user click accuracy errors, the interactive position usually does not completely coincide with any of the marked positions. To address this, this solution introduces a motion information estimation method based on relative position ratios. First, the two marked positions in the target image that are closest to the interactive position in the image synthesis direction (e.g., horizontal direction) are identified. Then, based on the coordinate distance between these two marked positions and the ratio of the interactive position relative to both, the relative position ratio f of the interactive position within that segment is calculated.
[0067] Therefore, based on the ratio f and the device motion information corresponding to the two marked positions, the target motion information corresponding to the interaction position can be derived.
[0068] Step 350: Based on the target motion information, control the target device to move to the target position and display the image captured by the target device at the target position.
[0069] The target location refers to the area that the user expects to view by performing a trigger operation in the target image; that is, the area that the user wants the target device to adjust its pointing or viewing angle for observation. Target motion information may include control parameters used to control the movement and / or turning of the target device to the target location, such as spatial coordinate information, viewing angle parameters, and lens orientation parameters.
[0070] Based on the target motion information determined by the aforementioned interaction positions, control commands for controlling the target device are further generated. These control commands may include a set of parameters used to drive the target device to perform corresponding spatial movements, such as the target viewing angle, rotation axis direction, horizontal / vertical deflection angle, and zoom magnification.
[0071] Furthermore, the control command can be sent to the target device via the Internet of Things, enabling the target device to perform motion control according to the direction or position indicated by the target motion information. For example, when the target device is a PTZ camera, the control command can guide the PTZ camera to rotate to the corresponding angle, so that the camera is aimed at the area desired by the user for observation.
[0072] Once the target device moves to the target location, it will acquire current image data at that location. The acquired image data is then sent back to the user terminal; after receiving the image data, the user terminal can display the captured image to the user.
[0073] Through the above process, a target image synthesized from multiple frames of original images captured by the target device is displayed, allowing users to directly perform click, touch, and other trigger operations on the target image. Based on the mapping relationship between the interaction position and each marked position and the motion information of the target device, the target motion information corresponding to the interaction position is derived, thereby controlling the target device to move to the target position. Through the above process of mapping interaction position, marked position, and motion information, the intuitive user interaction behavior can be transformed into target motion information that can be used to control the target device, avoiding cumbersome operations such as relying on angle input and preset paths. This solution improves the intuitiveness of device control, significantly improves the operating experience of the target device and the efficiency of user interaction, and solves the problems of cumbersome operation and inaccurate pointing in the user control process of PTZ camera devices.
[0074] In an exemplary embodiment, prior to step 310, the method may further include the following steps: acquiring multiple frames of original images captured by the target device under different motion information; and performing image synthesis processing on the multiple frames of original images to obtain the target image.
[0075] First, acquire multiple frames of raw images captured by the target device under different motion information. Motion information can include the target device's spatial coordinates, viewing angle parameters, and lens orientation parameters during image capture, representing the target device's position and orientation during image acquisition. In other words, different raw images correspond to image frames captured by the target device under different spatial positions, shooting angles, or lens orientations.
[0076] Therefore, image synthesis processing can be performed based on the acquired multiple frames of original images to generate the target image. Image synthesis processing can include steps such as image stitching, image registration, image fusion, projection transformation, and image enhancement, used to orderly integrate original images from multiple perspectives into a single target image with continuous content and complete structure. This target image can be displayed on the user terminal for user interaction and used for subsequent target device control operations.
[0077] Figure 4 A schematic diagram illustrating a specific implementation of acquiring raw images is shown, such as... Figure 4 As shown, the user terminal interface displays the real-time image captured by the target device. Below the user terminal interface, a "panoramic positioning" interactive interface area 401 is displayed. This area 401 is used to guide the user to generate and manage multiple original image acquisition angles and construct a panoramic image (i.e., the target image). Figure 5 These are three original images acquired by the target device. The horizontal progress bar on the user terminal interface and the status information "10% in panorama generation" displayed in the lower right corner indicate the current progress of the target image generation by the target device.
[0078] Figure 6 A schematic diagram illustrating a specific implementation of generating a target image is shown, such as... Figure 6 As shown, the user terminal interface displays the final generated target image. Additionally, the interface provides a "Regenerate Panoramic Image" button, which can be used to re-trigger the target image synthesis process after changes in the scene environment, updates to the original image, or manual adjustments to the viewing angle and acquisition path by the user. Based on the updated original image and its corresponding motion information, the system can re-stitch and fuse the images and re-mark the positions of each original image to generate a new target image, ensuring the timeliness and spatial coverage accuracy of the target image.
[0079] It should be noted that the system can control the target device to automatically perform movement operations according to a preset path and acquire corresponding image frames under different movement states. Specifically, based on preset control parameters such as horizontal angle step values and pitch angle ranges, the target device can sequentially rotate and capture multiple image frames covering a certain monitoring area. For example, the target device can perform image acquisition every 10 degrees of rotation until the entire monitoring area is covered. This automatic control method enables a highly efficient raw image acquisition process without human intervention.
[0080] In addition, raw image acquisition can also be performed through manual user control. For example, users can control the movement, rotation, pitch, or zoom of the target device using control buttons or swipe gestures on the user terminal interface. During the process of the user controlling the target device to switch perspectives and acquire raw images, the motion information of the target device corresponding to each frame of raw image (such as the spatial coordinates of the target device, perspective parameters, lens orientation parameters, etc.) can be recorded to facilitate subsequent image compositing processing.
[0081] Under the above embodiments, by acquiring multiple frames of original images collected by the target device under different motion information, and performing image synthesis processing on the multiple frames of original images to generate a target image, the target image has a more complete spatial coverage and higher visual continuity, which can provide users with an intuitive and unified interactive screen view. Thus, it can accurately capture the user's interactive behavior in the target image, and combined with the motion information of the original image, deduce the corresponding target motion information, so that the target device can move precisely to the area that the user wants to observe, significantly enhancing the interactivity and continuity of the target image, providing an accurate and convenient image basis for subsequent device control operations, and improving the overall interactive control experience and response accuracy based on the target image.
[0082] In an exemplary embodiment, the above method may further include the following steps: assigning corresponding marker positions to each original image based on motion information corresponding to multiple frames of original images; and generating a target image carrying each marker position according to the marker positions corresponding to each original image.
[0083] Specifically, while acquiring multiple frames of original images, motion information corresponding to the acquisition of each original image (such as spatial coordinate information of the target device, viewing angle parameters, lens orientation parameters, etc.) can also be acquired simultaneously to characterize the spatial location and shooting orientation covered by each original image.
[0084] Therefore, based on the spatial relationship and synthesis order between the original images, the synthesis position of each original image in the final target image can be determined. For example, if the original images are taken sequentially from left to right, the left view can be marked as the left side of the target image and the right view can be marked as the right side of the target image. Combined with the image center point, starting coordinates, or boundary coordinates, the marked position of that frame in the target image can be determined.
[0085] The marked location may include the coordinate information of the image region synthesized from the original image into the target image, or it may be used to identify the position of the image center point in the target image coordinate system, which is not limited here.
[0086] After obtaining the marker positions corresponding to each original image, image stitching and fusion operations can be performed based on these positions to generate a target image containing all the marker positions. This target image not only presents continuous images from multiple shooting angles but also achieves precise association with the target device's motion information through the recorded marker positions. Therefore, when a user performs a trigger operation at any position in the target image, the desired target motion information can be quickly and accurately deduced based on the relative relationship between the interaction position corresponding to the trigger operation and the marker positions, thereby controlling the target device to move to the target position.
[0087] Figures 7 to 8 A schematic diagram illustrating a specific implementation of an original image including marked locations and its synthesized target image is shown, such as... Figure 7 As shown, three original images captured by the target device under different motion conditions are displayed, labeled "1", "2", and "3" respectively. Each original image covers a portion of the scene; for example, position "1" mainly covers the left side of the living room, position "2" covers the central TV area, and position "3" covers the right side of the dining room aisle area. Figure 8 As shown, the synthesized target image continuously displays the scene areas covered by the original images "1", "2", and "3", achieving a panoramic restoration of the indoor scene while maintaining image consistency and visual naturalness. The positions of the markers "1", "2", and "3" in the target image are also preserved from those in the original images.
[0088] It should be noted that during the image synthesis process, although corresponding marker positions are assigned to each original image, and the original images are stitched and merged into the target image based on these marker positions, these marker positions are not visually presented in the final target image displayed to the user. In other words, the target image, as an image for user viewing and interaction, can be a complete picture after visual optimization processing, without containing any form of visual markers.
[0089] These marker positions exist only as intermediate parameters within the system to establish the mapping relationship between interactive positions and motion information, supporting subsequent user triggering operations and precise control of the target device. This non-explicit display method ensures both the accuracy of image synthesis and motion information alignment, and also considers the aesthetics and visual cleanliness of the user's operation, thereby further enhancing the human-computer interaction experience.
[0090] Figure 9 The diagram illustrates the specific implementation of the target image displayed on the user terminal interface, such as... Figure 9 As shown, the user terminal interface displays a target image generated by fusing multiple original images. This target image covers the entire scene captured by the target device and is presented in a panoramic view on the user interface. To improve the user's visual experience and interface cleanliness, the target image displayed on the user terminal does not explicitly show the corresponding markers in the original images. Users can interact with the target image by clicking anywhere on it; for example, ... Figure 9 As shown, the user completed the trigger operation by clicking on the door area located on the right side of the image, and the system determined the corresponding interaction location accordingly.
[0091] Under the above embodiments, marker positions can be assigned to each original image during the generation of the target image, and the original image can be synthesized in the target image based on the marker positions during image stitching and fusion operations. The marker positions serve as intermediate parameters, establishing a correspondence between the interactive position of the user performing the trigger operation in the target image and the motion information of the target device.
[0092] In an exemplary embodiment, the method may further include the following steps: obtaining at least one marker position carried by the target image; determining the marker position in the target image corresponding to the interaction position.
[0093] As mentioned earlier, the target image is generated in the image stitching and fusion operation based on multiple original images, and carries the marked positions corresponding to each original image. Therefore, during the process of displaying the target image on the user terminal, the system can detect the user's trigger operation on the target image, identify the interaction position corresponding to the trigger operation, and further determine the marked position corresponding to the interaction position in the marked positions.
[0094] Specifically, after detecting a user's triggering action in the target image, matching processing can be performed on multiple marked positions carried in the target image based on the interaction position corresponding to the triggering action, so as to determine the marked position corresponding to the interaction position.
[0095] In one possible implementation, the spatial distance between the interaction position and each marker position can be calculated first, and the marker position closest to the interaction position can be selected as the marker position corresponding to the interaction position; or, if the marker position records its coordinate information in the target image, it can be directly determined whether the interaction position falls within the area corresponding to a certain marker position. If so, then the marker position is determined to be the marker position corresponding to the interaction position.
[0096] Through the above process, based on the location markers carried by the target image, it is possible to determine which frame of the original image is most likely to correspond to the area of the interactive operation, and then deduce the motion information corresponding to that original image, thereby achieving a precise mapping between the user-triggered operation and the target motion information.
[0097] In an exemplary embodiment, step 350 may include the following steps: performing identification processing on the marker positions corresponding to each original image in the target image to obtain the position of each marker position in the target image.
[0098] Regarding recognition processing, it can refer to image recognition operations performed on the marked positions corresponding to each original image in the target image.
[0099] It should be noted that corresponding marker positions can be assigned to the original images during the image acquisition stage. For example, by adding numerical numbers (such as "1", "2", "3", etc.) to the original images, the marker positions carried in the original images can be retained in the target images during the image compositing process.
[0100] However, when original images are processed to generate target images (i.e., panoramic images) through image compositing, operations such as stitching alignment, projection transformation, and image cropping often cause some content of the original images to be shifted, scaled, or even lose information. This makes it difficult to directly reconstruct the accurate position coordinates of the marked locations contained in the original images in the target image. Especially when the target image is a wide-format panoramic image, the content of each original image is rearranged in visual space, and the absolute coordinates of its original marked locations are no longer directly usable.
[0101] Therefore, in order to accurately identify the marker positions corresponding to each original image in the target image, it is necessary to perform recognition processing operations, such as using optical character recognition (OCR) or deep learning model recognition, to extract the marker positions retained in the target image from the original image, and thereby determine the position of each marker position in the target image.
[0102] In one possible implementation, the recognition process can be based on optical character recognition (OCR) technology, which parses the marker positions (such as "1", "2", "3", etc.) contained in the target image to identify the original image, so as to obtain the position of each marker position in the target image.
[0103] In another possible implementation, the recognition process can use a trained convolutional neural network model to automatically identify and classify the marked positions (such as the numbers "1", "2", "3", etc.) in the target image corresponding to the original image.
[0104] Specifically, regions in the target image that may contain the marked locations can be extracted using a set sliding window or region candidate box. Each extracted region image is input into a trained convolutional neural network model, which extracts the image features of the region and outputs the probability distribution corresponding to each digit category. Based on the category with the highest probability, the digit label of the region image, i.e., the marked location, is determined.
[0105] It should be noted that although the marked positions may not be presented in a visible manner in the final target image displayed to the user (e.g., they may be hidden before display), the system can still accurately determine the position of each marked position in the target image based on the marked positions recorded before image synthesis or through OCR recognition, thereby realizing the mapping relationship between the subsequent interactive position and the target device motion information.
[0106] It should be noted that, in order to improve the accuracy of digital tag recognition, image preprocessing operations can be performed on the target image before performing recognition processing, and digital localization and segmentation operations can be performed when necessary.
[0107] In one possible implementation, image preprocessing operations may include grayscale conversion (e.g., converting a color image of the target image to a grayscale image to preserve brightness information), binarization (e.g., converting the target image to a black and white image to make the numbers and background clearer), denoising (e.g., removing small dots and interference lines), and / or tilt correction (e.g., keeping the numbers horizontal).
[0108] In another possible implementation, digital localization and segmentation can be performed by using techniques such as contour detection, connected component analysis, or bounding box regression to locate digital regions (i.e., the regions where the marked positions are located) one by one in the target image; the target image containing multiple digital regions is then segmented into multiple separate region images for subsequent recognition processing.
[0109] In the above process, even if the positions of each marker are not explicitly displayed in the target image, the location of each marker in the target image can still be determined based on the recognition process, ensuring the accuracy of human-computer interaction and the simplicity of the image display interface.
[0110] In an exemplary embodiment, the above method may further include the following steps: finding at least one adjacent marker position in the target image based on the interaction position; and obtaining a marker position whose distance meets a set condition based on the distance between each original image in the image synthesis direction, based on the at least one adjacent marker position and the interaction position, so as to determine the marker position in the target image corresponding to the interaction position.
[0111] In this process, one or more marker locations within a preset range of the interaction location can be found in the target image to determine at least one marker location adjacent to the interaction location.
[0112] Regarding the preset range, considering the density of the marker positions in the target image, if the marker positions are far apart, the preset area needs to be expanded to avoid omissions; if the marker positions are close together, the preset range can be reduced to avoid excessive computation.
[0113] It is understandable that, considering system response efficiency and real-time performance, the preset range should not be too large, otherwise it will significantly increase the amount of data processing and affect response latency. Therefore, the preset range can be dynamically adjusted in combination with the device performance to achieve a balance between search accuracy and processing overhead.
[0114] Therefore, after finding at least one marker position adjacent to the interaction position, the marker positions that are closer to the interaction position and meet the set conditions can be further selected based on the relative distance between the interaction position and each adjacent marker position in the image synthesis direction, and used as the final marker positions in the target image corresponding to the interaction position.
[0115] Regarding the setting conditions, they can be used to filter out the final marker position from multiple adjacent marker positions. Furthermore, the setting conditions can also include directional restrictions (i.e., in the image compositing direction). In this case, the setting conditions can be used to limit the maximum allowable distance between the interaction position and the marker position in the image compositing direction. For example, if the distance between the interaction position and a certain marker position in the image compositing direction exceeds a preset threshold, then the marker position is excluded to avoid misjudgment due to excessive distance.
[0116] Regarding distance calculation, both the interactive position and each marker position can be represented as coordinate points (x, y). Therefore, the distance between the interactive position and each marker position can be calculated separately, thus obtaining the distance between the interactive position and each marker position.
[0117] To further explain, in the image synthesis process, the target image is composed of multiple original images stitched together in a certain order and direction. For example, if the target device captures multiple viewpoints (such as left view, middle view, and right view) sequentially from left to right, the generated target image will present a visual effect of continuous unfolding from left to right in the horizontal direction; if the target device captures floor images layer by layer from bottom to top, the target image may be formed by stitching together in the vertical direction.
[0118] Therefore, in the process of fusing multiple original images to generate a target image, there can be an image merging direction, which guides the order and positional alignment of the original images during splicing and fusion. For example, the merging direction can be horizontal (e.g., left → right), vertical (e.g., below → above), or other custom spatial paths.
[0119] Because of the aforementioned image synthesis direction, when searching for marker positions related to interaction positions, distance should not be the sole criterion. Instead, the directionality of image stitching should be considered for restriction and sorting to improve the matching accuracy between marker positions and interaction positions.
[0120] For example, in a scenario where original images are stitched and merged in a left-to-right direction, the spatial relationship between the marked positions is mainly characterized by horizontal expansion. In this case, the horizontal distance (i.e., the difference in the horizontal coordinate) between the interactive position and each marked position should be calculated first, and this should be used as the main sorting basis.
[0121] For example, in a scenario where original images are stitched and merged in the top-to-bottom direction, the spatial relationship between marked positions is mainly characterized by vertical expansion. In this case, the vertical distance (i.e., the difference in the ordinate) should be calculated first for filtering.
[0122] In one possible implementation, the condition could be that the two marker positions closest to the interaction position in the image compositing direction are selected. Then, the two marker positions in the target image that have the minimum distance from the interaction position in the image compositing direction, such as the first marker position and the second marker position, can be determined as the final marker positions.
[0123] In the above process, the corresponding marker position in the target image can be accurately and quickly determined based on the user's interaction position. By filtering the marker positions that meet the set conditions in the image synthesis direction, misjudgment or deviation caused by interference from irrelevant marker positions can be effectively avoided, thereby further improving the control accuracy and response performance of the target device after receiving user input, and providing users with a more natural, efficient and accurate human-computer interaction experience.
[0124] In an exemplary embodiment, step 330 may further include the following steps: when there are corresponding first and second marker positions at the interaction position, based on the interaction position and the mapping relationship between each marker position and the motion information of the target device, find the first motion information corresponding to the first marker position and the second motion information corresponding to the second marker position; and determine the target motion information based on the interaction position, the first motion information and the second motion information.
[0125] As mentioned earlier, motion information may include the spatial coordinates, viewing angle parameters, and lens orientation parameters of the target device when acquiring the original image. Therefore, the corresponding first motion information and second motion information can be determined based on the mapping relationship between the first marker position, the second marker position and the motion information of the target device.
[0126] For example, if the first marker position is "2" and the second marker position is "3", the system can retrieve the corresponding spatial information, lens angle, or other relevant motion parameters from the mapping relationship between each marker position and the motion information of the target device, and record them as the first motion information and the second motion information, respectively. The first motion information and the second motion information can be used to characterize the motion state of the original image during the target device's shooting process, which helps to accurately map the interactive positions in the future.
[0127] Figure 10 The diagram illustrates a specific implementation of interactive and marker positions, such as... Figure 10 As shown, when a user clicks on the location of the door in the target image, this location becomes the interaction location. The first marker position corresponding to this interaction location is "2" and the second marker position is "3".
[0128] Therefore, based on the first motion information and the second motion information, the target motion information corresponding to the current interaction position can be further calculated according to the relative positional relationship between the first and second marker positions.
[0129] In the above process, when there are multiple marked positions, the existing mapping relationship can be effectively utilized and the interactive position can be combined to make calculations, thereby generating target motion information that is highly consistent with the semantics of user operation, significantly improving the device control accuracy and the naturalness of human-computer interaction driven by image interaction.
[0130] In an exemplary embodiment, the above method may further include the following steps: calculating the target motion information corresponding to the interaction position based on the relative position ratio between the first marker position and the second marker position, and based on the relative position ratio, the first position information, and the second position information.
[0131] The relative position ratio can refer to the proportion of the relative spatial position of the interaction position between the first and second marker positions.
[0132] Specifically, if the image compositing direction is horizontal (e.g., stitching from left to right), the relative position ratio can be expressed as the proportion of the interactive position in the horizontal direction relative to the first and second marker positions; if the image compositing direction is vertical (e.g., stitching from top to bottom), it can be expressed as the relative proportion in the vertical direction.
[0133] Regarding the calculation of relative position ratios, if the image is stitched together horizontally from left to right, the distance difference between the interactive position and the two marked positions in the horizontal direction can be compared to determine which marked position it is closer to, and thus determine its approximate proportional position in the entire interval. If the image is stitched together vertically from top to bottom, a similar distance comparison can be made in the vertical direction to obtain the relative position ratio.
[0134] For example, if an image is stitched together from left to right, and the first marker is located on the left and the second marker is located on the right, then the user's interaction position can also be mapped to a coordinate value in a certain horizontal direction. Then, the relative position ratio, that is, the horizontal offset of the user's interaction position relative to the first marker position, is the percentage of the total horizontal distance between the first and second marker positions.
[0135] For example, the relative position ratio f = (point.x - point1.x) / (point2.x - point1.x), where point.x represents the coordinates of the interactive position corresponding to the user's touch operation (such as the horizontal coordinates), and point1.x and point2.x represent the coordinates of the first and second marked positions, respectively. If f = 0, it means that the interactive position is exactly at the first marked position; if f = 1, it means that the interactive position is exactly at the second marked position; if f = 0.5, it means that the interactive position is at the midpoint between the two marked positions.
[0136] Based on the relative position ratio f, interpolation can be further used to calculate the target motion information corresponding to the interaction position according to the difference between the first position information and the second position information, thereby establishing a precise mapping relationship between the user operation position and the device motion space.
[0137] In one possible implementation, if the first motion information and the second motion information each contain multiple parameters (such as spatial coordinate information, viewpoint parameters, lens orientation parameters, etc.), each parameter can be weighted one by one. That is, the difference between the first motion information and the second motion information is calculated on each parameter according to the relative position ratio f, and finally the complete target motion information is obtained.
[0138] For example, the target motion information result.x = position1.x + (position2.x - position1.x) × f, where position1.x is the first motion information and position2.x is the second motion information.
[0139] Through the above process, given the first motion information corresponding to the first marker position and the second motion information corresponding to the second marker position, the target motion information corresponding to the interaction position can be calculated based on the relative position ratio. This realizes a mapping mechanism from image space to device motion space, providing accurate, continuous, and dynamic target motion information support for subsequent device control and other operations. This helps the system maintain consistency and real-time performance in complex interaction scenarios, thereby improving the user's responsiveness and naturalness of interaction.
[0140] In another exemplary embodiment, step 330 may further include the following steps: if there is a corresponding marker position at the interaction position, find the motion information corresponding to the marker position according to the interaction position and the mapping relationship between each marker position and the motion information of the target device; and use the motion information corresponding to the marker position as the target motion information corresponding to the interaction position.
[0141] Specifically, when only one marker position is found for the interaction position, or when only one adjacent marker position satisfies the set conditions in the image synthesis direction, the interaction position can be considered to have sufficient spatial proximity to the marker position. In this case, the system does not need to perform subsequent calculation steps and can directly extract the motion information carried by the marker position in the original image before image synthesis, such as spatial coordinate information, viewpoint parameters, and lens orientation parameters, and output it as the target motion information of the interaction position.
[0142] The target image generated in the aforementioned embodiment is composed of three original images stitched together in the image synthesis direction from left to right, corresponding to the left view, middle view and right view respectively. Each original image has a marked position (such as the numbers "1", "2" and "3"), and records the motion information of the target device at the corresponding shooting time (such as spatial coordinates, viewing angle parameters, lens orientation, etc.).
[0143] Suppose a user performs a click operation (i.e. trigger operation) in the target image, and the interaction position point falls in the left area of the corresponding left view. The system detects only one marker position point1 that meets the set conditions within the preset search range of point. This marker position is the position marked "1" in the left view, corresponding to the second frame in the original image.
[0144] At this point, since there is only one marker position and the interaction position that satisfy the preset conditions in the direction of image synthesis, the system can directly use the first motion information position1 associated with the marker position point1 (e.g., spatial coordinates (x1, y1, z1), viewing angle θ1, and lens orientation vector1) as the target motion information corresponding to the interaction position point, without the need for subsequent calculations.
[0145] In the above process, when the marked position is unique, it can efficiently complete the rapid mapping of the interaction position to the motion information of the target device, providing an instant device response basis for user interaction input and improving the overall real-time performance and accuracy of the system.
[0146] Figure 11 This is a schematic diagram illustrating the specific implementation of a device control method in an application scenario. In this scenario, the target device is a pan-tilt camera, and the raw images can be processed using software such as OpenCV and Stitching.
[0147] like Figure 11 As shown, firstly, the PTZ camera will perform multiple directional shots according to the preset shooting trajectory or angle command (such as rotating at certain angles) to acquire multiple raw images covering the target scene. For each raw image captured, the current PTZ camera's pitch angle and yaw angle, and other motion information (i.e., PTZ angle) must be recorded so that a mapping relationship between the marked position and the PTZ camera's motion information can be established later.
[0148] Specifically, the pan-tilt unit can be moved and rotated to take pictures using camera control commands, and the pictures can be stored in memory using data structures. For example, the image data structure can be: NSArray ×images = @[image1,image2,image3....] .
[0149] Simultaneously record the motion coordinates of the PTZ camera corresponding to each photo: NSArray× cameraPtzPosition=@[position1(x,y),position2(x,y),position3(x,y)…].
[0150] In this scenario, the user issues a command (assuming x:20, y:0) to control the pan-tilt camera to rotate. At this point, the camera displays an image. This means that if the user wants to view the current camera display by rotating the pan-tilt, they need to issue a command coordinate (assuming x:20, y:0) to the camera. This coordinate is the target coordinate where the camera needs to move to the target angle.
[0151] Furthermore, to assist in subsequent location recognition and mapping, each original image is numbered (e.g., 1, 2, 3...), that is, a marked location is added, and a number (e.g., a large number) is placed in a prominent position in the center of the original image to facilitate OCR recognition.
[0152] Then, software such as OpenCV and Stitching can be used to stitch images together, combining the original images with marked positions into the target image imageB.
[0153] Of course, for aesthetic purposes, the original image without the marked positions can also be stitched together using the same parameters to obtain the target image imageA, which can then be displayed to the user on the user's terminal for interaction.
[0154] For example, such as Figure 12 As shown, a corresponding panoramic image can be generated based on the parameters corresponding to image1, image2, and image3 in the image data structure. The generated panoramic image does not include the image's marked positions and is used for display on the user terminal interface.
[0155] Then, as Figure 13 As shown, based on the parameters corresponding to image1, image2, and image3 in the image data structure, each original image can be further edited, such as marking the center position of each original image with a number, and then using the edited images as parameters to obtain a panoramic image with the marked position, which is stored in memory for subsequent calculations. The corresponding numbers on each original image are mapped to the motion coordinates of the pan-tilt camera corresponding to each original image, cameraPtzPosition, to obtain the corresponding mapping table Map.
[0156] Users can then view the target image (imageA) on the user terminal interface, click on a point of interest (i.e., trigger an operation), and the clicked location becomes the interactive location.
[0157] After obtaining the interaction location, the system can first convert the image coordinates of that interaction location in imageA to the corresponding coordinates in imageB. Since imageA and imageB are generated based on the same parameters using an image stitching algorithm, they have a one-to-one correspondence in the image coordinate system. The system can directly use coordinate mapping to achieve the coordinate transformation from imageA to imageB.
[0158] Subsequently, the system uses OCR technology to identify the marked positions in imageB, and parses the marked positions (such as "1", "2", "3", etc.) in imageB that are used to identify the original image, in order to obtain the position of each marked position in the target image.
[0159] Specifically, at least one adjacent marker position can be found in imageB based on the interaction position; based on the at least one adjacent marker position and the interaction position, the marker position whose distance meets the set conditions is obtained according to the distance of each original image in the image synthesis direction, so as to determine the marker position in the target image corresponding to the interaction position.
[0160] When there are corresponding first and second marker positions at the interaction position, the first motion information corresponding to the first marker position and the second motion information corresponding to the second marker position are found according to the mapping relationship between the interaction position and the motion information of the PTZ camera. The relative position ratio between the interaction position and the first and second marker positions is used to calculate the target motion information corresponding to the interaction position.
[0161] For example, the relative position ratio f = (point.x - point1.x) / (point2.x - point1.x), where point.x represents the coordinates of the interactive position corresponding to the user's touch operation (such as the horizontal coordinate), and point1.x and point2.x represent the coordinates of the first and second marker positions, respectively. Then, the target motion information result.x = position1.x + (position2.x - position1.x) × f, where position1.x is the first motion information and position2.x is the second motion information.
[0162] If there is a corresponding marker position at the interaction position, the motion information corresponding to the marker position is found based on the mapping relationship between the interaction position, each marker position and the motion information of the PTZ camera; the motion information corresponding to the marker position is used as the target motion information corresponding to the interaction position.
[0163] Suppose a user performs a click operation (i.e., trigger operation) in the target image, and the interaction position point falls in the corresponding area of the middle view. The system detects only one marker position point2 that meets the set conditions within the preset search range of the point. This marker position is the position of the marker "2" in the middle view, corresponding to the second frame in the original image. Since there is only one marker position that meets the preset conditions in the image synthesis direction with the interaction position, the system can directly use the second motion information position2 associated with the marker position point2 (e.g., spatial coordinates (x2, y2, z2), viewing angle θ2, and lens orientation vector2) as the target motion information corresponding to the interaction position point, without the need for subsequent calculations.
[0164] Finally, based on the target motion information, the system generates a PTZ control command and sends it to the PTZ camera, driving the camera to move according to the target motion information. Once the camera reaches the target location, it can capture and display a real image of that location, achieving precise linkage between image interaction and PTZ control.
[0165] In this application scenario, by adding markers (e.g., embedding numbered labels at the center points of each image) to multiple original images captured by a PTZ camera, and then using image stitching technology to generate target images containing and without the markers, visual consistency between the two types of images is achieved. This method effectively improves the continuity of the interactive experience and avoids the problems of relying on manual input of rotation angles, complex operation, and low positioning accuracy in traditional preset position control methods.
[0166] Furthermore, a one-to-one mapping relationship is pre-established between each marked position and the motion information (such as spatial coordinate information, viewing angle parameters, lens orientation parameters, etc.) of the PTZ camera. This allows the system to calculate the target motion information corresponding to the interaction position based on the identified marked position and its corresponding motion information, combined with the spatial relationship between the user's interaction position and adjacent marked positions, when the user initiates a trigger operation (such as clicking or touching) at any position while viewing a target image without marked positions. This enables automated and high-precision positioning control.
[0167] This application significantly improves the automation and accuracy of motion information calculation, reducing the positioning error to within ±0.1°, which is far superior to traditional angle-based positioning methods (the error is generally above ±1°), avoiding the errors caused by manual operation or inaccurate angle presets in traditional methods.
[0168] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0169] The following are embodiments of the apparatus described in this application, which can be used to execute the device control method involved in this application. For details not disclosed in the embodiments of the apparatus described in this application, please refer to the method embodiments of the device control method involved in this application.
[0170] Please see Figure 14 This application provides a device control device 900, including but not limited to: an image display module 910, an information acquisition module 930, and a device control module 950.
[0171] The image display module 910 is used to display the target image; the target image is obtained based on multiple frames of original images captured by the target device.
[0172] The information acquisition module 930 is used to respond to a trigger operation on the target image, acquire the interaction position corresponding to the trigger operation and the marker position corresponding to each original image, so as to obtain the target motion information corresponding to the interaction position according to the mapping relationship between the interaction position and each marker position and the motion information of the target device.
[0173] The device control module 950 is used to control the target device to move to the target position based on the target motion information and display the image captured by the target device at the target position.
[0174] In one exemplary embodiment, the device control apparatus is further configured to acquire multiple frames of original images captured by the target device under different motion information; and to perform image synthesis processing on the multiple frames of original images to obtain the target image.
[0175] In an exemplary embodiment, the device control apparatus is further configured to assign corresponding marker positions to each original image based on motion information corresponding to multiple frames of original images; and generate a target image carrying each marker position according to the marker positions corresponding to each original image.
[0176] In one exemplary embodiment, the information acquisition module is further configured to acquire at least one marker position carried by the target image; and determine the marker position in the target image corresponding to the interaction position.
[0177] In an exemplary embodiment, the information acquisition module is further configured to identify the marker positions corresponding to each original image in the target image, and obtain the position of each marker position in the target image.
[0178] In an exemplary embodiment, the information acquisition module is further configured to find at least one adjacent marker position in the target image based on the interaction position; and based on the at least one adjacent marker position and the interaction position, obtain the marker position whose distance meets the set conditions according to the distance between each original image in the image synthesis direction, so as to determine the marker position in the target image corresponding to the interaction position.
[0179] In an exemplary embodiment, the information acquisition module is further configured to, when there are corresponding first and second marker positions at the interaction position, search for first motion information corresponding to the first marker position and second motion information corresponding to the second marker position based on the interaction position and the mapping relationship between each marker position and the motion information of the target device; and determine the target motion information based on the interaction position, the first motion information and the second motion information.
[0180] In an exemplary embodiment, the information acquisition module is further configured to calculate target motion information corresponding to the interaction position based on the relative position ratio between the first marker position and the second marker position, as well as the first position information and the second position information.
[0181] In an exemplary embodiment, the information acquisition module is further configured to, when there is a corresponding marker position at the interaction position, search for motion information corresponding to the marker position based on the interaction position and the mapping relationship between each marker position and the motion information of the target device; and use the motion information corresponding to the marker position as the target motion information corresponding to the interaction position.
[0182] It should be noted that the device control device provided in the above embodiments is only illustrated by the division of the above functional modules when controlling the device. In actual applications, the above functions can be assigned to different functional modules as needed. That is, the internal structure of the device control device will be divided into different functional modules to complete all or part of the functions described above.
[0183] Furthermore, the device control apparatus and device control method embodiments provided in the above embodiments belong to the same concept, and the specific way in which each module performs operations has been described in detail in the method embodiments, and will not be repeated here.
[0184] Please see Figure 15 This application provides an electronic device 4000, which may include smartphones, tablets, etc.
[0185] exist Figure 15 In this context, the electronic device 4000 includes at least one processor 4001 and at least one memory 4003.
[0186] Data interaction between the processor 4001 and the memory 4003 can be achieved through at least one communication bus 4002. This communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used to represent it in the figure, but this does not indicate that there is only one bus or one type of bus.
[0187] Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.
[0188] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0189] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device capable of storing static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device capable of storing information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing computer programs having instruction or data structure forms and accessible by electronic device 400, but not limited to these.
[0190] The memory 4003 stores a computer program, and the processor 4001 can read the computer program stored in the memory 4003 through the communication bus 4002.
[0191] The computer program is executed by one or more processors 4001 to implement the device control methods in the above embodiments.
[0192] Furthermore, this application provides a storage medium storing a computer program, which is executed by one or more processors to implement the device control method described above.
[0193] This application provides a computer program product, including a computer program that is executed by one or more processors to implement the device control method described above.
[0194] Compared with related technologies, this solution displays a target image synthesized from multiple frames of original images captured by the target device. This allows users to directly perform clicks, touches, and other trigger operations on the target image. Based on the mapping relationship between the interaction position and each marked position and the motion information of the target device, the target motion information corresponding to the interaction position is derived, thereby controlling the target device to move to the target position. Through the above process of mapping interaction position, marked position, and motion information, the intuitive user interaction behavior can be transformed into target motion information that can be used to control the target device, avoiding cumbersome operations such as relying on angle input and preset paths. This solution improves the intuitiveness of device control, significantly improves the operating experience of the target device and the efficiency of user interaction, and solves the problems of cumbersome operation and inaccurate positioning of PTZ camera devices during user control.
[0195] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A device control method, characterized in that, include: Display the target image; The target image is obtained based on multiple frames of original images acquired by the target device; In response to a trigger operation on the target image, the interaction position corresponding to the trigger operation and the marker position corresponding to each of the original images are obtained, so as to obtain the target motion information corresponding to the interaction position according to the mapping relationship between the interaction position and each of the marker positions and the motion information of the target device; Based on the target motion information, the target device is controlled to move to the target position and the image captured by the target device at the target position is displayed.
2. The method as described in claim 1, characterized in that, Before displaying the target image, the method includes: Acquire multiple frames of raw images of the target device under different motion conditions; The target image is obtained by performing image synthesis processing on the multiple original images.
3. The method as described in claim 2, characterized in that, Before performing image synthesis processing on the multiple original images to obtain the target image, the method further includes: Based on the motion information corresponding to the multiple original images, a corresponding marker position is assigned to each of the original images; Based on the marker positions corresponding to each of the original images, a target image carrying each of the marker positions is generated.
4. The method as described in claim 1, characterized in that, The step of responding to a trigger operation on the target image and obtaining the interaction position corresponding to the trigger operation and the marker position corresponding to each of the original images includes: Obtain at least one marker location carried by the target image; Determine the marker position in the target image corresponding to the interaction position.
5. The method as described in claim 4, characterized in that, The step of obtaining at least one marker location carried by the target image includes: The marker positions corresponding to each of the original images in the target image are identified to obtain the position of each marker position in the target image.
6. The method as described in claim 4, characterized in that, Determining the marker position in the target image corresponding to the interaction position includes: Based on the interaction location, at least one adjacent marker location is found in the target image; Based on the at least one adjacent marker position and the interaction position, and according to the distance between each of the original images in the image synthesis direction, a marker position whose distance meets a set condition is obtained, so as to determine the marker position in the target image corresponding to the interaction position.
7. The method as described in claim 1, characterized in that, The step of obtaining the target motion information corresponding to the interaction position based on the interaction position and the mapping relationship between each of the marked positions and the motion information of the target device includes: If there are corresponding first and second marker positions at the interaction position, the first motion information corresponding to the first marker position and the second motion information corresponding to the second marker position are found according to the interaction position and the mapping relationship between each marker position and the motion information of the target device. The target motion information is determined based on the interaction location, the first motion information, and the second motion information.
8. The method as described in claim 7, characterized in that, Determining the target motion information based on the interaction location, the first motion information, and the second motion information includes: Based on the relative position ratio of the interaction position between the first marker position and the second marker position; Based on the relative position ratio, the first position information, and the second position information, the target motion information corresponding to the interaction position is calculated.
9. The method as described in claim 1, characterized in that, The step of obtaining the target motion information corresponding to the interaction position based on the interaction position and the mapping relationship between each of the marked positions and the motion information of the target device includes: If a corresponding marker position exists at the interaction position, the motion information corresponding to the marker position is found based on the interaction position and the mapping relationship between each marker position and the motion information of the target device. The motion information corresponding to the marked position is used as the target motion information corresponding to the interaction position.
10. A device control apparatus, characterized in that, include: The image display module is used to display the target image; The target image is obtained based on multiple frames of original images acquired by the target device; The information acquisition module is used to respond to a trigger operation on the target image, acquire the interaction position corresponding to the trigger operation and the marker position corresponding to each of the original images, so as to obtain the target motion information corresponding to the interaction position according to the mapping relationship between the interaction position and each of the marker positions and the motion information of the target device; The device control module is used to control the target device to move to the target position based on the target motion information and display the image captured by the target device at the target position.
11. An electronic device comprising at least one processor and at least one memory, wherein, The memory stores a computer program, characterized in that, when the computer program is executed by the processor, it implements the device control method as described in any one of claims 1 to 9.
12. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by one or more processors, it implements the device control method as described in any one of claims 1 to 9.
13. A computer program product comprising a computer program, characterized in that, When the computer program is executed by one or more processors, it implements the device control method as described in any one of claims 1 to 9.