Image processing method, device control method, apparatus, device, and storage medium

CN122824980APending Publication Date: 2026-09-25SHENZHEN LUMIUNITED TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610820507.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本申请提供了一种图像处理方法、设备控制方法、装置、电子设备、存储介质及计算机程序产品,可以解决相关技术中存在的用户在查看目标的一些细节时操作过于繁琐且操作效率低下的问题

Benefits of technology

在上述技术方案中,用户终端首先向用户展示第一画面,该第一画面是目标设备针对目标空间进行拍摄和采集得到的画面,若用户不满足该第一画面,便可以借助用户终端在第一画面上重新确定画面区域,此时,用户终端便能够检测到在第一画面为确定画面区域而触发的操作,以此获取到用于指示用户重新确定的画面区域在第一画面中的位置的第一位置数据,并基于该第一位置数据通知目标设备进行摄像角度和/或焦距的调整,最终向用户展示第二画面,该第二画面是针对目标设备针对空间区域进行拍摄和采集得到的画面,该空间区域即是由用户重新确定的画面区域在目标空间中映射确定的,在整个图像处理过程中,用户仅进行了简单操作,便能够查看到符合其心意的第二画面,有效地避免了用户在查看目标的一些细节时需要与用户终端之间进行多次交互,从而能够有效地解决相关技术中存在的用户在查看目标的一些细节时操作过于繁琐且操作效率低下的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824980A_ABST
    Figure CN122824980A_ABST
Patent Text Reader

Abstract

The application provides an image processing method and device, electronic equipment and storage medium, and relates to the technical field of Internet of Things. The image processing method comprises the following steps: displaying a first picture; the first picture is a picture obtained by photographing and collecting a target space by a target device; if an operation triggered when the first picture is a certain picture area is detected, first position data is acquired; the first position data is used to indicate the position of the picture area in the first picture; a second picture obtained based on the first position data is displayed; the second picture is a picture obtained by photographing and collecting a space area by the target device after the photographing angle and / or focal length of the target device are adjusted based on second position data; and the second position data is mapped from the first position data and is used to indicate the position of the space area in the target space under the field of view angle of the target device. The application solves the problems of complicated operation and low operation efficiency of a user when checking some details of a target in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet of Things (IoT) technology, and more specifically, to an image processing method, a device control method, an apparatus, a device, and a storage medium. Background Technology

[0002] Currently, if users are not satisfied with the footage captured by the camera, they can remotely control the camera via their terminal to adjust the camera angle and focus so that the adjusted camera can capture and collect footage that satisfies the user.

[0003] However, in some scenarios, such as viewing house numbers, license plate numbers, and product labels, users need to first control the camera to point at the target using their user terminal, and then control the camera to zoom in on the target to view details. The entire process may require multiple interactions between the user and the user terminal, which is not only too cumbersome but also inefficient, seriously affecting the user experience. Summary of the Invention

[0004] This application provides an image processing method, a device control method, an apparatus, an electronic device, a storage medium, and a computer program product, which can solve the problems of overly cumbersome and inefficient operation for users when viewing details of a target in related technologies. The technical solution provided by this application is as follows: According to one aspect of this application, an image processing method is provided, the method comprising: displaying a first screen; the first screen being a screen captured and acquired by a target device for a target space; if an operation triggered in the first screen for determining a screen area is detected, acquiring first position data; the first position data being used to indicate the position of the screen area in the first screen; displaying a second screen acquired based on the first position data; the second screen being a screen captured and acquired by the target device for a spatial area after adjusting the camera angle and / or focal length based on the second position data; the second position data being mapped from the first position data and used to indicate the position of the spatial area in the target space under the field of view of the target device.

[0005] According to one aspect of this application, a device control method includes: receiving first position data; the first position data indicating the position of a screen region in a first screen; the screen region being determined in response to an operation triggered in the first screen to determine the screen region; the first screen being a screen captured and acquired by a target device targeting a target space; mapping the first position data from the coordinate system of the first screen to the coordinate system of the target device to obtain second position data; the second position data indicating the position of a spatial region in the target space under the field of view of the target device; and controlling the target device to adjust the camera angle and / or focal length based on the second position data, so that the target device, after adjusting the camera angle and / or focal length, captures and acquires a second screen targeting the spatial region.

[0006] According to one aspect of this application, an image processing apparatus includes: a first display module for displaying a first image; the first image being an image captured and acquired by a target device targeting a target space; a first data acquisition module for acquiring first position data if an operation triggered in the first image to determine an image area is detected; the first position data indicating the position of the image area in the first image; and a second display module for displaying a second image acquired based on the first position data; the second image being an image captured and acquired by the target device targeting a spatial area after adjusting the camera angle and / or focal length based on the second position data; the second position data being mapped from the first position data and used to indicate the position of the spatial area in the target space under the field of view of the target device.

[0007] According to one aspect of this application, a device control apparatus includes: a second data acquisition module for receiving first position data; the first position data indicating the position of a screen area in a first screen; the screen area being determined in response to an operation triggered in the first screen to determine the screen area; the first screen being a screen captured and acquired by a target device targeting a target space; a data mapping module for mapping the first position data from the coordinate system of the first screen to the coordinate system of the target device to obtain second position data; the second position data indicating the position of a spatial area in the target space under the field of view of the target device; and a device control module for controlling the target device to adjust the camera angle and / or focal length based on the second position data, so that the target device, after adjusting the camera angle and / or focal length, captures and acquires a second screen targeting the spatial area.

[0008] According to one aspect of this application, an electronic device includes at least one processor and at least one memory, wherein the memory stores a computer program that, when executed by the processor, implements the image processing method as described above.

[0009] According to one aspect of this application, a storage medium having a computer program stored thereon, which, when executed by one or more processors, implements the image processing method as described above.

[0010] According to one aspect of this application, a computer program product includes a computer program that, when executed by one or more processors, implements the image processing method as described above.

[0011] The beneficial effects of the above-mentioned technical solution provided in this application are: In the above technical solution, the user terminal first displays a first image to the user. This first image is an image captured and collected by the target device in relation to the target space. If the user is not satisfied with the first image, they can use the user terminal to redefine the image area on the first image. At this time, the user terminal can detect the operation triggered in the first image to redefine the image area, thereby obtaining first position data to indicate the position of the redefined image area in the first image. Based on the first position data, the user terminal notifies the target device to adjust the camera angle and / or focal length, and finally displays a second image to the user. This second image is an image captured and collected by the target device in relation to the spatial area, which is determined by mapping the redefined image area in the target space. In the entire image processing process, the user only needs to perform a simple operation to view the second image that meets their needs. This effectively avoids the need for multiple interactions between the user and the user terminal when viewing some details of the target, thus effectively solving the problem of overly cumbersome and inefficient operation when viewing some details of the target in related technologies. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a schematic diagram based on the implementation environment involved in this application; Figure 2 This is a hardware structure diagram of an electronic device according to an exemplary embodiment; Figure 3 This is a flowchart illustrating an image processing method according to an exemplary embodiment; Figure 4 This is a schematic diagram illustrating the adjustment from a first screen to a second screen according to an exemplary embodiment; Figure 5 yes Figure 3 A flowchart of step 320 in one embodiment corresponds to the following example; Figure 6 This is a schematic diagram illustrating various controls for implementing area positioning procedures and device positioning procedures according to an exemplary embodiment; Figure 7 This is a flowchart illustrating another image processing method according to an exemplary embodiment; Figure 8 This is a flowchart illustrating a device control method according to an exemplary embodiment; Figure 9 yes Figure 8 A flowchart of step 520 in one embodiment corresponds to the following example; Figures 10a to 10c This is a schematic diagram illustrating the specific implementation of an image processing method in an application scenario; Figure 11 This is a structural block diagram of an image processing apparatus according to an exemplary embodiment; Figure 12 This is a structural block diagram of a device control apparatus according to an exemplary embodiment; Figure 13 This is a structural block diagram of an electronic device according to an exemplary embodiment. Detailed Implementation

[0014] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0015] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this disclosure means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0016] A pan-tilt-zoom (PTZ) camera is a smart camera that can be remotely controlled to rotate horizontally (pan), vertically (tilt), and zoom. Users typically operate the PTZ camera through a control application on their mobile device (phone or tablet) to flexibly adjust the monitored view.

[0017] Currently, the control methods for PTZ cameras include the following: First, the directional key control method. Users control the gimbal rotation using the up, down, left, and right directional keys, requiring multiple clicks to move the image to the target position. This control method involves a long operation path, requiring users to repeatedly try to adjust the target area to near the center of the screen, and it cannot precisely control the range of movement.

[0018] Second, the touch positioning method. The user taps a location on the live screen, and the gimbal automatically rotates to center that location on the screen. This method solves the gimbal turning problem, but after positioning, the user still needs to zoom in to view details; positioning and zooming are two independent processes.

[0019] Third, the zoom control method. Users adjust the digital zoom level using a separate zoom button (+ / -) or slider, usually offering fixed levels (such as 1x, 2x, 4x, 8x) or continuous adjustment. Users need to rely on experience to judge which level to select, often requiring multiple attempts to find the appropriate zoom level.

[0020] Fourth, the gesture zoom method. Users zoom in and out on the screen by pinching with two fingers, but the zoom center is usually fixed in the current center of the screen, and the position of the area of ​​interest cannot be adjusted at the same time. If the target is not in the center of the screen after zooming, the positioning operation needs to be performed again.

[0021] In real-world usage scenarios, users often need to complete the entire process of locating a target location and zooming in to view details, such as house numbers, facial features, license plate numbers, and product labels. However, the above control methods still have the following defects and shortcomings in practical use cases: First, the positioning and zooming operations are separated, making the process cumbersome. Users need to first use the directional keys or touch positioning to rotate the gimbal to the target position, and then zoom in using a separate zoom button or gesture. Completing a full positioning and zooming process usually requires multiple steps, such as: clicking touch positioning, waiting for the screen to update, clicking the zoom button, waiting for the screen to update, fine-tuning the position, and adjusting the zoom again. The entire process may take tens of seconds or even longer, resulting in low operational efficiency.

[0022] Secondly, the magnification accuracy is difficult to control, placing a heavy cognitive burden on users. Existing digital zoom typically uses fixed zoom levels or manual sliders, making it difficult for users to determine the appropriate magnification to clearly see the entire target area. For example, a user might want to view details within a rectangular area of ​​an image but doesn't know the optimal zoom level for that area, requiring repeated attempts at different zoom levels, increasing operational complexity and cognitive burden.

[0023] Secondly, repeated adjustments are inefficient and significantly impacted by network latency. Due to network latency, users must wait for server response and screen updates for each operation when using the directional keys for positioning and zooming. The accumulated network latency during the repeated adjustments of positioning, zooming, repositioning, and re-zooming severely affects the user experience, especially in environments with poor network conditions, where the round-trip latency for each interaction can reach hundreds of milliseconds or even several seconds.

[0024] Finally, the operational logic is abstract and doesn't match the user's intuitive needs. The user's intuitive need is "I want to see the content of this rectangular area in the image clearly," but the current method requires the user to understand and operate two separate sets of operational logic for gimbal rotation control and digital zoom control. The user needs to break down the intuitive intention of "selecting an area" into two independent steps: "first rotate to the center of the area, then zoom in to the appropriate magnification." This indirect operation path increases the learning cost.

[0025] To address the aforementioned problems, this application provides an image processing method applicable to an image processing device. This image processing device can be deployed on an electronic device, which may refer to an electronic device with display and control functions, such as a smartphone, tablet computer, laptop computer, desktop computer, smart control panel, etc.; the electronic device may also be an electronic device with central control functions, such as a gateway, server, etc.

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0027] Figure 1 This is a schematic diagram of an implementation environment for an image processing method. The implementation environment includes at least a user terminal 110, a smart device 130, a server 170, and network equipment. Figure 1 In this context, network devices include gateway 150 and router 190, but this is not intended to be a specific limitation.

[0028] The user terminal 110, which can also be considered as a user terminal, terminal, or front end, can deploy (or install) the client associated with the smart device 130. This user terminal 110 can be an electronic device such as a smartphone, tablet, laptop, desktop computer, smart control panel, or other device with display and control functions, without limitation.

[0029] The client, associated with the smart device 130, is essentially where the user registers an account and configures the smart device 130. For example, the configuration includes adding a device identifier to the smart device 130, so that when the client runs on the user terminal 110, it can provide the user with functions such as device display and device control of the smart device 130. This client can be in the form of an application or a web page. Correspondingly, the interface for displaying the device on the client can be in the form of a program window or a web page, and there is no limitation here.

[0030] Smart device 130 is deployed in gateway 150 and communicates with gateway 150 through its own configured communication module, thereby being controlled by gateway 150. It should be understood that smart device 130 generally refers to one of multiple smart devices 130. This application embodiment only uses smart device 130 as an example; that is, this application embodiment does not limit the number or type of smart devices deployed in gateway 150. In one application scenario, smart device 130 is deployed in gateway 150 by accessing it through a local area network. The process of smart device 130 accessing gateway 150 through a local area network includes: gateway 150 first establishes a local area network, and smart device 130 joins the local area network established by gateway 150 by connecting to it. This local area network includes, but is not limited to, ZIGBEE or Bluetooth. Among them, the smart device 130 can be a smart printer, a smart fax machine, a smart camera (such as a PTZ camera), a smart air conditioner, a smart door lock, a smart light, or a human body sensor, door and window sensor, temperature and humidity sensor, water immersion sensor, natural gas alarm, smoke alarm, wall switch, wall socket, wireless switch, wireless wall sticker switch, cube controller, curtain motor, millimeter wave radar, etc., equipped with a communication module.

[0031] The interaction between user terminal 110 and smart device 130 can be achieved through a local area network (LAN) or a wide area network (WAN). In one application scenario, user terminal 110 establishes a wired or wireless communication connection with gateway 150 via router 190, such as Wi-Fi, allowing user terminal 110 and gateway 150 to be deployed on the same LAN, thus enabling user terminal 110 to interact with smart device 130 via the LAN path. In another application scenario, user terminal 110 establishes a wired or wireless communication connection with gateway 150 via server 170, such as 2G, 3G, 4G, 5G, or Wi-Fi, allowing user terminal 110 and gateway 150 to be deployed on the same WAN, thus enabling user terminal 110 to interact with smart device 130 via the WAN path.

[0032] The server-side 170 can also be considered as the cloud, cloud platform, platform side, or backend, etc. This server-side 170 can be a single server, a server cluster consisting of multiple servers, or a cloud computing center consisting of multiple servers, in order to better provide backend services to a massive number of user terminals 110. For example, backend services include device control services.

[0033] In one application scenario, the user terminal 110 first displays a first screen to the user. This first screen is the image captured and collected by the target device 130 (such as a PTZ camera) in the target space. If the user is not satisfied with the first screen, they can use the user terminal 110 to redefine the screen area on the first screen. Accordingly, the user terminal 110 can detect the operation triggered to redefine the screen area on the first screen, thereby obtaining the first position data. This first position data is used to indicate the position of the screen area redefine by the user in the first screen.

[0034] For the target device 130, the camera angle and / or focal length will be adjusted based on the second location data mapped from the first location data. This second location data indicates the position of the spatial region within the target space from the perspective of the target device 130. The spatial region will then be photographed and captured to obtain a second image. It should be noted that in this application scenario, the user terminal 110 first sends the first location data to the gateway 150 or the server 170. The gateway 150 or server 170 then completes the mapping process from the first location data to the second location data. Subsequently, the gateway 150 transmits the second location data to the target device 130 via a local area network, or the server 170 transmits the second location data to the target device 130 via a wide area network. Of course, in other application scenarios, the user terminal 110 can also directly send the first location data to the target device 130, so that the mapping process from the first location data to the second location data can also be completed independently by the target device 130. This is not a specific limitation.

[0035] Through the interaction between the user terminal 110 and the target device 130, the user terminal 110 ultimately displays a second screen to the user.

[0036] Throughout the entire process of the target device 130 switching from shooting at the target space to shooting at the spatial area, the user only needs to perform a simple operation to view the second image that meets their expectations. This effectively avoids the need for multiple interactions between the user and the user terminal when viewing some details of the target, thus effectively solving the problem of overly cumbersome and inefficient operation when viewing some details of the target in related technologies.

[0037] Please see Figure 2 , Figure 2 This is a hardware structure diagram of an electronic device according to an exemplary embodiment. This electronic device is suitable for... Figure 1 The user terminal 110, gateway 150, or server 170 in the implementation environment are shown.

[0038] It should be noted that this electronic device is merely an example adapted to this application and should not be construed as providing any limitation on the scope of use of this application. Furthermore, this electronic device should not be interpreted as requiring or depending on any specific feature. Figure 2 One or more components of the exemplary electronic device 200 shown.

[0039] The hardware structure of electronic device 200 can vary significantly due to differences in configuration or performance, such as... Figure 2 As shown, the electronic device 200 includes: a power supply 210, an interface 230, at least one memory 250, and at least one central processing unit (CPU) 270.

[0040] Specifically, power supply 210 is used to provide operating voltage for various hardware devices on electronic device 200.

[0041] Interface 230 includes at least one wired or wireless network interface 231 for interacting with external devices. For example, to perform... Figure 1 The diagram illustrates the interaction between user terminal 110 and smart device 130 in the implementation environment.

[0042] Of course, in other examples adapted in this application, interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input / output interface 235, and at least one USB interface 237, etc. Figure 2 As shown, this does not constitute a specific limitation.

[0043] The memory 250 serves as a carrier for resource storage and can be a read-only memory, random access memory, disk, or optical disk, etc. The resources stored on it include the operating system 251, application programs 253, and data 255, etc., and the storage method can be temporary storage or permanent storage.

[0044] The operating system 251 is used to manage and control the various hardware devices and application programs 253 on the electronic device 200, so as to enable the central processing unit 270 to perform calculations and processing on the massive data 255 in the memory 250. It can be Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0045] Application 253 is a computer program formed by computer-readable instructions based on operating system 251 to perform at least one specific task, and may include at least one module ( Figure 2 (Not shown), each module may contain corresponding computer-readable instructions. For example, the image processing device can be considered as an application 253 deployed on electronic device 200.

[0046] Data 255 can be static images (such as the first screen and the second screen), dynamic images (such as videos) stored on the disk, or first position data, second position data, etc., stored in memory 250.

[0047] The central processing unit 270 may include one or more processors and is configured to communicate with the memory 250 via at least one communication bus to read computer programs stored in the memory 250, thereby performing operations and processing on massive amounts of data 255 stored in the memory 250. For example, an image processing method or a device control method may be implemented by the central processing unit 270 reading an application program 253 stored in the memory 250.

[0048] Furthermore, this application can also be implemented through hardware circuits or a combination of hardware circuits and software. Therefore, the implementation of this application is not limited to any specific hardware circuit, software, or combination thereof.

[0049] Please see Figure 3 This application provides an image processing method applicable to electronic devices, such as electronic devices that can be... Figure 1 The user terminal 110 shown in the implementation environment can have the following hardware structure: Figure 2 As shown.

[0050] In the following method embodiments, for ease of description, the user terminal is used as the execution subject of each step of the method, but this does not constitute a specific limitation.

[0051] like Figure 3 As shown, the method may include the following steps: Step 310: Display the first screen.

[0052] The first image is the picture captured and recorded by the target device in the target space. The target device can be... Figure 1 The diagram shows an intelligent device with image capture and acquisition capabilities in the implementation environment; for example, this intelligent device could be a PTZ camera. The target space refers to the physical space currently being monitored by the target device; for example, the target space could be an indoor room, an outdoor courtyard, an underground parking lot, etc.

[0053] It is understood that shooting can be a single shot or continuous shooting. For continuous shooting, a video clip can be obtained, and the first frame can be any frame from that video. Conversely, multiple shots can result in multiple photographs, and the first frame can be any one of those photographs. In other words, the first frame in this embodiment can come from a moving image, such as any frame from a video clip, or from a static image, such as any one of multiple photographs. Accordingly, image processing in this embodiment can be performed frame by frame.

[0054] Regarding the acquisition of the first frame, it can originate from images captured and collected in real time by the target device, or it can be images captured and collected by the target device within a historical time period pre-stored on the user terminal. Therefore, for the user terminal, after the target device captures and collects the first frame, it can process the first frame in real time, or it can pre-store the first frame for further processing, for example, displaying and processing the first frame according to user instructions. Thus, the image processing in this embodiment can be applied to the first frame acquired in real time or to the first frame acquired within a historical time period; no limitation is imposed here.

[0055] Step 320: If an operation triggered to determine a screen area is detected in the first screen, the first position data is obtained.

[0056] The first position data is used to indicate the position of the screen area in the first screen. Since the screen area is determined based on the user's actions on the first screen, the first position data can also be considered to be related to the user's actions on the first screen to determine the screen area. Alternatively, it can be understood that the first position data corresponds to the user's actions on the first screen. If the user's actions on the first screen are different, the position of the screen area in the first screen will be different, and correspondingly, the first position data will also be different.

[0057] It should be understood that if a user is not satisfied with the first view of the target space displayed on the user terminal, for example, if some details of the target space seen by the user in the first view are not clear enough, the user will hope that the target device can re-capture and re-capture the corresponding images for some details of the target space. Accordingly, the image area refers to the area that the user re-selects in the first view for some details of the target space.

[0058] Figure 4 A diagram illustrating the positional relationship between the first frame and the frame area is shown. Figure 4In this context, the first screen 301 includes a screen area. The first screen 301 actually displays the complete target space to the user, while the screen area displays some details of the target space to the user. It can also be understood that, based on the screen area redefined by the user on the first screen, a portion of the target space will be displayed to the user.

[0059] In this embodiment, the user redefines the screen area based on some details in the target space in the first screen, which is achieved based on operation triggering. For example... Figure 4 As shown, a user can draw a rectangle (or a closed curve of other shapes such as an ellipse or circle) on the first screen. The user terminal can then detect the rectangle and obtain the screen area determined by the user on the first screen. That is, the rectangle is regarded as the screen area determined by the user on the first screen. Alternatively, a user can draw a straight line on the first screen. The user terminal can then detect the straight line and obtain the screen area determined by the user on the first screen. That is, the straight line is regarded as the diagonal of the screen area determined by the user on the first screen. Therefore, the above drawing operations can all be regarded as operations triggered on the first screen to determine the screen area.

[0060] It should be noted that the form of operation can vary depending on the input components configured on the user terminal. For example, if the user terminal is a smartphone and the configured input component is a touch screen, the operation can be a gesture operation such as clicking or swiping; if the user terminal is a desktop computer and the configured input component is a mouse, the operation can be a mechanical operation such as clicking, double-clicking, or dragging.

[0061] After the user terminal obtains the screen area determined by the user on the first screen, it can determine the position of the screen area in the first screen, thereby obtaining the first position data used to indicate the position of the screen area in the first screen.

[0062] In some embodiments, the first position data may include coordinate data and / or size data of the screen area. In some embodiments, the coordinate data of the screen area may include the start and end coordinates of the screen area's diagonal. In other embodiments, the coordinate data of the screen area may also include the center coordinates of the screen area. In some embodiments, the size data of the screen area may include the width and height of the screen area. In other embodiments, the size data of the screen area may also include the aspect ratio of the area.

[0063] Step 330: Display the second screen based on the data obtained from the first location.

[0064] The second image is the image captured and collected by the target device after adjusting the camera angle and / or focal length based on the second position data, targeting a spatial area. The spatial area refers to the physical space corresponding to the image area in the first image; it can also be understood as the physical space currently monitored by the target device after adjusting the camera angle and / or focal length.

[0065] So, the spatial area refers to some details of the target space that the user wants to see (that is, a part of the target space). Accordingly, unlike the first screen which shows the user the entire target space, the second screen shows the user a part of the target space.

[0066] First, it should be noted that the second position data is obtained by mapping the first position data and is used to indicate the position of the spatial region in the target space under the field of view of the target device.

[0067] As mentioned earlier, the first position data is used to indicate the position of the screen area within the first frame, in order to Figure 1 In the illustrated implementation environment, after obtaining the first location data, the user terminal 110 sends the first location data to the gateway 150, requesting the gateway 150 to control the target device 130 to adjust the camera angle and / or focal length. This allows the target device 130 to capture and collect images of the corresponding spatial area within the target space, thus satisfying the user's desire to view a second image of that spatial area. Therefore, the gateway 150 needs to map the position of the image area in the first image to the position of the spatial area within the target space under the target device's field of view. In other words, it needs to map the first location data from the coordinate system of the first image to the second location data in the coordinate system of the target device to achieve the adjustment of the target device's camera angle and / or focal length.

[0068] Secondly, the second location data may include the camera angle and / or focal length of the target device to be adjusted.

[0069] In some embodiments, the camera angle to be adjusted for the target device is calculated based on the coordinate data of the image area in the first position data. In some embodiments, the focal length to be adjusted for the target device is calculated based on the size data of the image area in the first position data.

[0070] Continue with Figure 1 As shown in the example implementation environment, after determining the camera angle and / or focal length to be adjusted of the target device 130, the gateway 150 can generate the corresponding device control command and send the device control command to the target device 130, so that the target device 130 responds to the device control command to perform the corresponding rotation action (i.e., adjust the camera angle) and / or zoom action (i.e., adjust the focal length).

[0071] Therefore, for user terminal 110, after target device 130 adjusts the camera angle and / or focal length, it can receive and display the second image captured and collected by target device 130 for the spatial area.

[0072] Of course, in other embodiments, the above mapping process can also be completed independently by the target device, or it can be understood that the image processing method is implemented through the interaction between the user terminal and the target device, which does not constitute a specific limitation.

[0073] Through the above process, a complete interactive loop is achieved, allowing users to move from "seeing the entire target space" to "determining the image area" and then to "seeing a portion of the target space" with simple operations. The entire image processing process reduces user operations from 6 to 10 steps to 1 step, and shortens operation time from 10 to 15 seconds to 2 to 3 seconds, improving efficiency by more than 5 times. This effectively avoids the need for multiple interactions between the user and the user terminal when viewing details of the target, significantly reducing the user's operational burden and cognitive load. Thus, it effectively solves the problem of overly cumbersome and inefficient operations when viewing details of the target in related technologies.

[0074] As mentioned earlier, depending on the input components configured on the user terminal, the operations triggered to define the screen area on the first screen can have different forms of expression. For example, there are touch screen-based gesture operations such as clicking, swiping, and long-pressing. Based on this, the following will combine... Figure 5 Taking the touch screen as an example of the input component configured on the user terminal, the method for determining the screen area through the recognition of the first gesture operation is explained in detail: Please see Figure 5 In one exemplary embodiment, step 320 may include the following steps: Step 321: Based on the point of action on the first screen, determine whether the first gesture operation is detected on the first screen.

[0075] The first gesture operation represents the user's gesture operation to select a screen area in the first screen, and is used to indicate the selection of a screen area in the first screen.

[0076] First, it should be noted that the point of contact refers to the position where a user's finger or stylus touches the touchscreen; it can also be considered a touch point or control point. For the user terminal, detecting the point of contact on the first screen can identify whether the operation triggered on the first screen is a first gesture operation.

[0077] In some embodiments, the first gesture operation may be a gesture operation triggered by the user drawing a closed curve (such as a rectangle, ellipse, circle, etc.) in the first screen. Accordingly, the first gesture operation of selecting a screen area in the first screen may be a continuous touch action in which the point of action slides from the starting position and eventually returns to the starting position (that is, the ending position).

[0078] In other embodiments, the first gesture operation can also be a gesture operation triggered by the user drawing a straight line in the first screen. Accordingly, the first gesture operation of selecting a screen area in the first screen can be manifested as a continuous touch action in which the point of action slides from the starting position to the ending position.

[0079] It should be noted that the starting position can be considered as the touch position corresponding to the point of action when the user's finger or stylus is pressed down, and the ending position can be considered as the touch position corresponding to the point of action when the user's finger or stylus is lifted up.

[0080] In some embodiments, the user terminal can determine whether a first gesture operation is detected on the first screen using a touch event listener. For example, the detection process of the first gesture operation may include the following steps: when an action point is detected on the first screen, the touch event listener acquires the touch event sequence corresponding to the action point; if the touch event sequence includes a press event, a move event, and a release event in sequence, then it is determined that a first gesture operation has been detected. The press event instructs the user's finger or stylus to press down on the touch screen; the move event instructs the user's gesture or stylus to move on the touch screen; and the release event instructs the user's finger or stylus to release on the touch screen.

[0081] In other embodiments, the user terminal can also determine whether a first gesture operation has been detected on the first screen by the length of the movement trajectory of the point of action. Specifically, the detection process of the first gesture operation may include the following steps: when a point of action is detected on the first screen, the movement trajectory of the point of action on the first screen is tracked; if the movement trajectory exceeds a distance threshold, it is determined that a first gesture operation has been detected.

[0082] The movement trajectory refers to the continuous sequence of all positions traversed by the point of action on the first screen from its starting position to its ending position.

[0083] It is understood that minor movements of a user's finger or stylus on the touchscreen (such as the natural tremor when a user's finger accidentally touches the touchscreen) should not be interpreted as the user intending to reselect a screen area on the initial screen. In other words, the user terminal can only determine that a first gesture operation has been detected when the movement trajectory exceeds a distance threshold. This distance threshold can be flexibly adjusted according to the actual needs of the application scenario and is not limited here.

[0084] In this approach, noise can be effectively filtered out by using a distance threshold, ensuring that the user terminal can respond promptly to the first gesture operation that reflects the user's intention to select a screen area, thus avoiding the frequent triggering of unnecessary screen area selection processes.

[0085] Of course, in other embodiments, the criteria for determining the first gesture operation are not limited to whether the length of the movement trajectory exceeds a distance threshold. They can also be combined with parameters such as the movement speed corresponding to the point of application, the pressure applied, or the touch area for a comprehensive judgment. For example, if the movement speed is too fast (exceeding a speed threshold), it is determined that the first gesture operation was not detected.

[0086] If the user terminal detects the first gesture operation, indicating that the user intends to redefine the screen area, that is, the user is redefining the screen area on the first screen, then step 322 is executed.

[0087] Conversely, if the user terminal does not detect the first gesture operation, it indicates that the user has not redefined the screen area on the first screen. In this case, the process can return to step 321 and continue to determine whether the first gesture operation has been detected on the first screen based on the point of action on the first screen.

[0088] In addition, if the user terminal detects a point of action on the first screen but not the first gesture operation, it indicates that the user has other intentions. For example, the other intention may be the intention to manually locate, that is, the user does not intend to view further details of the target space, but only wants to manually control the target device to slightly adjust the camera angle and / or focal length.

[0089] At this point, the user terminal can display a third screen to the user. This third screen is the image captured and collected by the target device after positioning, targeting the target space. Specifically, the process of realizing the manual positioning intention may include the following steps: if no operation triggered to determine the screen area is detected on the first screen, the target device is instructed to position the focus based on the position of the point of action when it is on the first screen; then the third screen is displayed.

[0090] In this context, the focal point refers to the point at which the target device can achieve a clear image. Essentially, the area in the image that is in focus is the clearest. Therefore, by using focal point positioning, the camera angle of the target device will be adjusted to ensure that the target device is centered within the image and that the corresponding area is in the clearest view.

[0091] For example, such as Figure 6As shown, if the user clicks on the location of the screen area on the first screen, the user terminal can detect the point of action at that location on the first screen. Since the user does not perform any other operations afterward, the user terminal cannot detect the first gesture operation or other gesture operations after detecting the point of action. At this time, the user terminal determines that the user has the intention to manually locate, and will notify the target device to adjust the focus to the location so that the target device can re-capture and collect the target space based on the adjusted focus.

[0092] by Figure 1 Taking the implementation environment shown as an example, the user terminal 110 sends the coordinate data of its location to the gateway 150. The gateway 150 first calculates the camera angle to be adjusted of the target device 130 based on the coordinate data of the location, and generates a corresponding device control command and sends it to the target device 130. Accordingly, the target device 130 can receive the device control command and respond to the device control command to perform the corresponding rotation action, that is, adjust the camera angle.

[0093] Therefore, for user terminal 110, after target device 130 adjusts its camera angle, it can receive and display the third image re-captured and collected by target device 130 for the target space, and continue reading. Figure 6 It is worth mentioning that, unlike the second screen, the third screen still shows the entire target space rather than a part of it. Also, although both screens show the entire target space, the third screen is different from the first screen. The target device captures the third screen and the first screen under different focus conditions.

[0094] This approach enables intelligent switching between regional positioning and focal positioning. Users no longer need to manually switch between the two positioning modes. The user terminal can automatically determine the user's intent based on the point of action detection and execute the operation corresponding to the user's intent. This covers a wider range of application scenarios and improves the adaptability of the entire image processing workflow.

[0095] Step 322: Based on the start and end positions of the action point of the first gesture operation on the first screen, determine the screen area in the first screen.

[0096] As mentioned earlier, the starting position can be considered as the touch position corresponding to the point of action when the user's finger or stylus is pressed down, and the ending position can be considered as the touch position corresponding to the point of action when the user's finger or stylus is lifted up.

[0097] In some embodiments, the position of the screen area on the first screen is determined by using the start position and end position as a set of diagonal vertices of the screen area defined by the user on the first screen. For example, the coordinates of the start position on the first screen can be represented as (x1, y1) and the coordinates of the end position on the first screen can be represented as (x2, y2). By storing these coordinates, the screen area can be determined in the first screen.

[0098] Of course, in other embodiments, the determination of the screen area may not be based on a set of diagonal vertices of the screen area, but on the circle and radius of the screen area. For example, the user first clicks a position (x0, y0) as the circle of the screen area, and then slides a straight line outward a distance r as the radius of the screen area. Starting from this starting position, the user slides and eventually returns to the starting position to form a circle. Finally, the user terminal can determine the screen area in the first screen by storing these coordinates.

[0099] Step 323: Obtain first position data based on the position of the screen area in the first screen.

[0100] The first position data is used to indicate the position of the screen area in the first screen.

[0101] In some embodiments, the first position data may include coordinate data and / or size data of the screen area. In some embodiments, the coordinate data of the screen area may include the start and end coordinates of the screen area's diagonal. In other embodiments, the coordinate data of the screen area may also include the center coordinates of the screen area. In some embodiments, the size data of the screen area may include the width and height of the screen area. In other embodiments, the size data of the screen area may also include the aspect ratio of the area.

[0102] Taking the aforementioned screen area as a rectangle as an example, the first position data may include the starting position coordinates (x1, y1), the ending position coordinates (x2, y2), and the center point coordinates (x_center, y_center) of the screen area as coordinate data, and may also include the width w_select, height h_select, and area aspect ratio ratio_select of the screen area as size data.

[0103] Under the above embodiments, a standardized processing pipeline is established from detecting the first gesture operation to outputting the first position data. Through this pipeline, the user terminal can accurately identify the user's true intention and other intentions regarding redefining the screen area, providing data basis for subsequently accurately controlling the target device to adjust the camera angle and / or focal length to reshoot and acquire the second screen.

[0104] In an exemplary embodiment, prior to step 320, the method may further include the following steps: Step 410: Determine whether the region selection function is enabled.

[0105] The area selection function is used to indicate whether it is allowed to select an area of ​​the screen in the first screen.

[0106] The inventors recognized that users can perform both a primary gesture (i.e., a single-finger swipe) and other gestures on the first screen. These other gestures include, for example, a single-finger tap to manually locate the target device, and a two-finger pinch to zoom in and out of the first screen. Furthermore, due to the type of gesture, even the primary gesture may correspond to different user intentions. Therefore, this embodiment includes a region selection function to prevent conflicts between primary gestures corresponding to different user intentions or between primary gestures and other gestures.

[0107] In some embodiments, the user terminal can maintain a state variable to record the on / off state of the region selection function. Then, each time the user terminal detects a hovering point on the first screen, it first checks this state variable to determine whether the region selection function is enabled based on the indication of the state variable. For example, a state variable of 1 indicates that the region selection function is enabled, and a state variable of 0 indicates that the region selection function is disabled.

[0108] When the area selection function is enabled, it means that the user terminal allows the user to select an area on the first screen, thus allowing step 320 to be executed. At this time, the user terminal can determine that the user wants to reselect an area on the first screen upon detecting the first gesture operation.

[0109] Conversely, when the area selection function is not enabled, it means that the user terminal does not allow the user to select a screen area in the first screen, and step 320 is prohibited. At this time, even if the user terminal detects the first gesture operation, it is regarded as the user having other intentions rather than the intention to reselect a screen area in the first screen.

[0110] Based on this, by introducing a function switch to realize different interaction modes, the first gesture operation in different interaction modes does not interfere with other gesture operations, thereby effectively avoiding the limitation of the limited number of gestures, and thus effectively preventing conflicts between the first gesture operations corresponding to different user intentions or between the first gesture operation and other gesture operations.

[0111] In some embodiments, the region selection function can be activated based on a second gesture. This second gesture differs from the first gesture and represents the user's activation of the region selection function. This second gesture can be a long press, a double tap, etc., and is not limited here. Taking a long press as an example, the process of activating the region selection function may include the following steps: On the first screen, detect whether the dwell time of the point of action at the current position exceeds a time threshold; if so, mark the region selection function as activated. Here, the dwell time refers to the time interval between the point of action at the current position from the moment it is pressed to the current moment.

[0112] The user terminal can start a timer to count from the moment the touch point is detected pressing at the current position. If the touch point moves before the timer's count exceeds the time threshold T_press, the timer stops counting, and the second gesture is determined not to be a long press. If the timer's count exceeds the time threshold T_press and the touch point remains at the current position, the second gesture is determined to be a long press, and the region selection function is enabled. The time threshold T_press can be flexibly adjusted according to the actual needs of the application scenario and is not limited here.

[0113] In some embodiments, enabling the region selection function can also be achieved based on a region determination control provided by the user terminal. Specifically, the process of enabling the region selection function may include the following steps: determining whether an operation triggered by switching the region determination control from an inactive state to an active state is detected; if so, marking the region determination control as active to indicate that the selection of a screen region in the first screen is allowed.

[0114] In other words, the user terminal can display a region selection control to the user, which can indicate whether selecting a first region on the first screen is allowed through different states. Specifically, the active region selection control indicates that selecting a first region on the first screen is permitted. See also... Figure 4 , Figure 4 This demonstrates the different states of the area definition control, such as Figure 4 As shown, black icon 302 represents an active area determination control, and white icon 305 represents an inactive area determination control.

[0115] Therefore, for a region selection control that is inactive, when a user clicks the region selection control, the user terminal can detect this click operation and know that the user wants to reselect a region in the first screen. At this time, the user terminal can respond to the click operation by switching the region selection control from an inactive state to an active state, thereby enabling the region selection function. Here, the click operation is considered to be the operation triggered by switching the region selection control from an inactive state to an active state.

[0116] As mentioned earlier, the region selection function can be enabled based on a second gesture operation. However, since this second gesture operation may not directly affect the region selection control, it may cause the region selection control to remain inactive. Consequently, the inactive region selection control may still indicate that the selection of a screen area is not allowed in the first screen, which conflicts with the already enabled region selection function. In this case, the user terminal needs to switch the region selection control from an inactive state to an active state.

[0117] In some embodiments, when the area selection function is enabled, the user terminal can mark the area selection control as active, thereby indicating that the selection of a screen area in the first screen is allowed.

[0118] In other embodiments, if the user terminal does not detect any operation triggered by the region determination control (such as the click operation mentioned above), or does not detect the point of action remaining on the first screen, or detects that the point of action remaining on the first screen does not stay at the current position for more than a time threshold, the user terminal's region selection function is not enabled. In this case, the user terminal can mark the region determination control as inactive to indicate that selecting a screen area in the first screen is not allowed.

[0119] It should be noted that the switching between different states of the area definition control can be represented in at least one of the following ways: the color of the area definition control changes from gray to a highlight color (such as green or blue); the icon shape of the area definition control changes (such as from a hollow rectangle to a solid rectangle); the area definition control changes from having no effect to having a breathing light effect; the area definition control changes from having no text label to displaying a text label (such as "On"), and no limitation is imposed here.

[0120] With the cooperation of the above embodiments, a switch management mechanism for the region selection function is introduced, and two different user interaction modes, long press activation and switch activation, are realized. This allows the first gesture operation under different user intentions to coexist with other gesture operations without interference. This greatly enhances the flexibility of gesture operation when the types of gesture operation are limited, thereby greatly enriching the user experience.

[0121] Please see Figure 7In an exemplary embodiment, after step 310, the method may further include the following steps: Step 340: If an operation triggered by adjusting the camera angle and / or focal length of the target device is detected, then the device adjustment data is acquired.

[0122] Among them, the equipment adjustment data is used to indicate the camera angle and / or focal length to be adjusted on the target equipment.

[0123] The inventors realized that in real-world applications, users do not always need to perform the entire process of re-regioning the screen (i.e., region positioning); sometimes, only minor adjustments to the target device's camera angle and / or focal length are required. Therefore, in this embodiment, An alternative device positioning process is provided, making the frame area positioning process compatible and coexisting with the device positioning process.

[0124] In some embodiments, the user terminal may provide an input dialog box to allow the user to input the camera angle and / or focal length to be adjusted for the target device.

[0125] In other embodiments, the user terminal provides device adjustment controls for adjusting the camera angle and / or focal length of the target device.

[0126] Please refer back to Figure 4 ,exist Figure 4 In the process, the device adjustment controls may include an angle adjustment control 304 and a focal length adjustment control 303. The angle adjustment control 304 is used to adjust the camera angle of the target device, and the focal length adjustment control 303 is used to adjust the focal length of the target device.

[0127] like Figure 4 As shown, on the one hand, users can click any of the directional keys (such as up, down, left, right, etc.) displayed by the angle adjustment control 304 to adjust the camera angle of the target device. On the other hand, users can also click any of the icons displayed by the focus adjustment control 303; different icons correspond to different zoom levels, thereby adjusting the focus of the target device. Figure 4 The icon clicked by the user is displayed in black. Accordingly, the user's terminal can detect the click operation and obtain the corresponding device adjustment data. This click operation can be considered as an operation triggered by adjusting the camera angle and / or focal length of the target device.

[0128] In some embodiments, the region positioning process and the device positioning process are configured to be mutually exclusive. That is, if the region determination control is active, indicating that the region selection function is enabled, the user terminal is allowed to execute the region positioning process, which allows the user to reselect a screen region on the first screen. If the region determination control is inactive, indicating that the region selection function is disabled, the user terminal is allowed to execute the device positioning process, that is, the user terminal detects whether there is an operation triggered by adjusting the camera angle and / or focus for the target device.

[0129] For example, first, the system determines whether the region selection function is enabled, i.e., whether the region determination control is active. If the region determination control is active, the user terminal is allowed to execute the region positioning process. When an operation triggered by the region determination control is detected, the region determination control is switched from active to inactive. Next, when the region determination control is inactive, the user terminal can detect whether there is an operation triggered by adjusting the camera angle and / or focal length of the target device, thereby obtaining device adjustment data. Then, the user terminal can notify the target device to adjust the camera angle and / or focal length based on the device adjustment data, so that the target device can re-capture and record the target space based on the adjusted camera angle and / or focal length.

[0130] by Figure 1 Taking the implementation environment shown as an example, the user terminal 110 sends the device adjustment data to the gateway 150. The gateway 150 first determines the camera angle and / or focal length to be adjusted of the target device 130 based on the device adjustment data, and generates a corresponding device control command and sends it to the target device 130. Accordingly, the target device 130 can receive the device control command and respond to the device control command to perform the corresponding rotation and / or scaling actions.

[0131] Therefore, for user terminal 110, after target device 130 adjusts the camera angle and / or focal length, it can receive the fourth image that the target device 130 re-captures and collects for the target space, and then execute step 350.

[0132] Step 350 displays the fourth screen based on the data acquisition adjusted by the second device.

[0133] The fourth image is the image captured and collected by the target device after adjusting the camera angle and / or focal length based on the data adjusted by the second device, targeting the target space.

[0134] It is worth mentioning that, unlike the second screen, the fourth screen shows the entire target space rather than a part of it. Also, although both screens show the entire target space, the fourth screen is different from the first screen. The target device is shown in the first and fourth screens at different camera angles and / or focal lengths.

[0135] Under the above embodiments, users can flexibly choose between "regional positioning" or "device positioning" according to the actual needs of the application scenario, covering a wider range of application scenarios and improving the adaptability of the entire image processing process.

[0136] Please see Figure 8 This application provides a device control method, which is applicable to electronic devices. For example, the electronic device may be... Figure 1 The gateway 150 shown in the implementation environment can also be... Figure 1 The server-side 170 in the implementation environment is shown. The hardware structure of this electronic device can be as follows: Figure 2 As shown.

[0137] In the following method embodiments, for ease of description, the gateway is used as the execution subject of each step of the method, but this does not constitute a specific limitation.

[0138] like Figure 8 As shown, the method may include the following steps: Step 510: Receive the first position data.

[0139] The first position data indicates the location of the image area within the first frame. The image area is determined in response to an operation triggered in the first frame to define the image area. The first frame is the image captured and acquired by the target device targeting the target space.

[0140] by Figure 1 In the example implementation environment shown, after obtaining the first location data, the user terminal 110 sends the first location data to the gateway 150. Correspondingly, the gateway 150 can receive the first location data and control the target device 130 to adjust the camera angle and / or focal length so that the target device 130 can shoot and collect images of the corresponding spatial area in the target space, thereby satisfying the user's desire to view a second image of the spatial area.

[0141] Step 520: Map the first position data from the coordinate system of the first screen to the coordinate system of the target device to obtain the second position data.

[0142] The second position data is used to indicate the position of the spatial region in the target space under the field of view of the target device.

[0143] First, it should be noted that the coordinate system of the first screen refers to the pixel coordinate system (in pixels) on the user terminal. The coordinate system of the target device is the device coordinate system of the target device, which can also be understood as the world coordinate system (in degrees or in terms of the target's position in space) of the target device.

[0144] Because there is a fundamental difference between the pixel coordinate system on the user terminal and the coordinate system of the target device—they represent different dimensions in the space where the user terminal and the target device reside—the first position data sent by the user terminal cannot be directly used to control the target device without coordinate system transformation. Therefore, in this embodiment, before adjusting the camera angle and / or focal length of the target device, a mapping between the first and second position data is required.

[0145] Secondly, it should be noted that the second location data includes the camera angle and / or focal length of the target device to be adjusted.

[0146] In some embodiments, the camera angle to be adjusted on the target device is calculated based on coordinate data in the first position data. This coordinate data includes, but is not limited to: the starting and ending position coordinates on the diagonal of the screen area, and the center position coordinates of the screen area.

[0147] In some embodiments, the focal length to be adjusted for the target device is calculated based on dimensional data in the first position data. This dimensional data includes, but is not limited to, the width, height, and aspect ratio of the image area.

[0148] Step 530: Based on the second position data, control the target device to adjust the camera angle and / or focal length so that the target device can capture and obtain a second image of the spatial area after adjusting the camera angle and / or focal length.

[0149] For the gateway, the camera angle and / or focal length to be adjusted of the target device in the second location data are first encapsulated into a device control command. As the gateway interacts with the target device, the device control command is sent to the target device so that the target device responds to the device control command and performs the corresponding rotation and / or scaling actions.

[0150] In the above process, the user only needs to perform a simple operation, and the target device can automatically adjust the camera angle and / or focal length so that the user can see a second image that meets their expectations. This effectively avoids the need for multiple interactions between the user and the user terminal when viewing some details of the target, thus effectively solving the problem of overly cumbersome and inefficient operation when viewing some details of the target in related technologies.

[0151] Please see Figure 9 In one exemplary embodiment, step 520 may include the following steps: Step 521: Based on the first position data, calculate the camera angle and / or focal length to be adjusted for the target device.

[0152] Among them, the parameters corresponding to the camera angle can include the horizontal rotation angle alpha and the vertical rotation angle beta, and the parameters corresponding to the focal length can include the zoom factor.

[0153] In some embodiments, the parameter calculation process may include an angle calculation subprocess and a focal length calculation subprocess.

[0154] In some embodiments, the angle calculation subprocess may include the following steps: determining the center position of the screen area based on the first position data; calculating the deviation between the center position of the screen area and the center position of the first screen; and calculating the camera angle to be adjusted of the target device based on the deviation and the field of view of the target device.

[0155] For example, based on the starting position coordinates (x1, y1) and ending position coordinates (x2, y2) of the screen area, calculate the center position coordinates (x_center, y_center), where x_center = (x1+x2) / 2, y_center = (y1+y2) / 2.

[0156] The center position of the first frame is located at coordinates (W / 2, H / 2), where W is the width of the first frame and H is the height. Therefore, the deviation between the two can be calculated: the horizontal offset delta_x = x_center - W / 2, and the vertical offset delta_y = y_center - H / 2. The offset is in pixels; a positive value indicates a rightward or downward offset, and a negative value indicates a leftward or upward offset. The sign of the deviation directly corresponds to the direction of rotation of the target device. For example, a positive delta_x indicates that the target device needs to rotate to the right, and a negative delta_x indicates that the target device needs to rotate to the left; a positive delta_y indicates that the target device needs to rotate downward, and a negative delta_y indicates that the target device needs to rotate upward.

[0157] The field of view (FOV) of the target device refers to the maximum range of spatial angles that the target device can capture, including the horizontal field of view (FOV_h) and the vertical field of view (FOV_v). The field of view is determined by the focal length and image sensor size of the target device. For example, the horizontal field of view (FOV_h) can be between 60 and 90 degrees, and the vertical field of view (FOV_v) can be between 45 and 60 degrees.

[0158] Accordingly, the formula for calculating the horizontal rotation angle alpha is alpha = (delta_x / W) × FOV_h, and the formula for calculating the vertical rotation angle beta is beta = (delta_y / H) × FOV_v.

[0159] Assuming the width W of the first frame is 1920 pixels, the horizontal field of view FOV_h is 60 degrees, and the horizontal offset delta_x is 480 pixels, then the horizontal rotation angle alpha = (480 / 1920) × 60 = 15 degrees, which means that the target device needs to rotate 15 degrees to the right.

[0160] In some embodiments, the focal length calculation sub-process may include the following steps: determining the width, height, and aspect ratio of the image area based on first position data; the aspect ratio refers to the ratio of the width to the height of the image area; if the aspect ratio is less than the image aspect ratio, then calculating the focal length to be adjusted for the target device based on the height of the first image and the height of the image area; the image aspect ratio refers to the ratio of the width to the height of the first image; if the aspect ratio is greater than or equal to the image aspect ratio, then calculating the focal length to be adjusted for the target device based on the width of the first image and the width of the image area.

[0161] For example, the aspect ratio (ratio_select) refers to the ratio of the width (w_select) to the height (h_select) of the screen area, i.e., ratio_select = w_select / h_select.

[0162] The aspect ratio (ratio_screen) refers to the ratio of the width (W) to the height (H) of the first screen element, i.e., ratio_screen = W / H. The screen width (W) and height (H) can be dynamically adapted to the resolution of the user's touchscreen.

[0163] If the area's aspect ratio `ratio_select` is less than the screen's aspect ratio `ratio_screen`, it means the screen area is narrower and taller than the first screen. In this case, the height of the first screen is a limiting factor. If the screen is zoomed in based on its width, the zoomed-in second screen may not be able to fully accommodate the screen area in the height direction. In this case, the scaling factor is calculated based on the height `H` of the first screen and the height `h_select` of the screen area, using the formula: `zoom1 = H / h_select`.

[0164] If the area's aspect ratio `ratio_select` is greater than or equal to the screen's aspect ratio `ratio_screen`, it means the screen area is wider and flatter than the first screen. In this case, the width of the first screen is the limiting factor. If the screen is zoomed in based on its height, the zoomed-in second screen may not be able to completely accommodate the screen area in its width direction. In this situation, the scaling factor is calculated based on the width `W` of the first screen and the width `w_select` of the screen area, using the formula: `zoom2 = W / w_select`.

[0165] As an alternative, the focal length calculation subprocess can also avoid the binary choice strategy based on width or height. Instead, it can employ a smooth transition strategy, that is, weighted fusion of the width-based scaling factor and the height-based scaling factor based on the proximity of ratio_select and ratio_screen. For example, when ratio_select and ratio_screen are very close (e.g., the difference is less than 0.1), the average of the scaling factors zoom1 and zoom2 in both directions is taken, making the transition from the first frame to the second frame more natural.

[0166] In this approach, the angle calculation sub-process and the focal length calculation sub-process can be performed simultaneously by two independent threads or processing units, reducing processing latency and thus effectively improving equipment control efficiency.

[0167] Step 522: Obtain the second position data from the camera angle and / or focal length to be adjusted of the target device.

[0168] The second location data includes the camera angle and / or focal length of the target device to be adjusted.

[0169] With the cooperation of the above embodiments, the mapping from the first location data to the second location data is realized, so that the first location data sent by the user terminal can be applied to the control of the target device, thereby enabling the image processing flow from the first screen to the second screen to be realized.

[0170] The inventors realized that the adjustment range of target devices is usually limited. Taking a pan-tilt camera as an example, the camera angle of a pan-tilt camera has physical limits (e.g., from 0 degrees to 355 degrees horizontally, and from -90 degrees to 90 degrees vertically). Similarly, the focal length is also limited by the upper limit of the maximum optical / digital zoom capability. If the parameters corresponding to the calculated camera angle and / or focal length are not limited, it is very likely to lead to unpredictable results such as damage to the target device or failure to execute device control commands.

[0171] In an exemplary embodiment, the process of limiting the camera angle and / or focal length to be adjusted for the target device may include the following steps: Step 610: If the camera angle and / or focal length to be adjusted of the target device exceeds the adjustment range of the target device, then the camera angle and / or focal length to be adjusted of the target device are restricted based on the adjustment range of the target device.

[0172] In some embodiments, the limiting process may employ a clamping operation, which specifically means that if the camera angle and / or focal length to be adjusted of the target device is greater than the maximum value of the adjustment range, then the camera angle and / or focal length to be adjusted of the target device is set to the maximum value; if the camera angle and / or focal length to be adjusted of the target device is less than the minimum value of the adjustment range, then the camera angle and / or focal length to be adjusted of the target device is set to the minimum value; otherwise, the camera angle and / or focal length to be adjusted of the target device remains unchanged.

[0173] In other embodiments, the limiting process may also include a smoothing operation. Specifically, when the gateway detects that the camera angle and / or focal length to be adjusted of the target device frequently touches the upper and lower limits of the adjustment range, it may also adopt a soft clamping strategy to gradually reduce the adjustment range of the camera angle and / or focal length to be adjusted of the target device near the upper and lower limits of the adjustment range, so as to avoid the target device from oscillating repeatedly at the upper and lower limits of the adjustment range.

[0174] Step 630: Send the second notification message.

[0175] The second notification message is used to indicate that the target device's adjusted camera angle and / or focal length have reached the adjustment range.

[0176] It is understandable that if the gateway directly restricts the camera angle and / or focal length to be adjusted on the target device without notifying the user, the user may be confused as to why the second screen still does not fully display the area they want to view. Therefore, in this embodiment, the gateway can send a second prompt message to the user terminal to promptly inform the user that the adjustment of the target device has reached its limit, avoiding repeated, ineffective attempts by the user.

[0177] Based on the above embodiments, the target device's amplitude limiting protection and user prompt mechanism are realized, which fully ensures the balance between device control automation and device physical constraints. This not only effectively prevents damage to the target device, but also avoids user confusion caused by restrictions on the adjustment of the target device, greatly improving the user experience.

[0178] Figures 10a to 10c This is a schematic diagram illustrating the specific implementation of an image processing method in an application scenario. Figure 10a This document presents a flowchart illustrating an image processing method used in this application scenario. Figure 10b The diagram illustrates the "long press to activate + swipe" interaction mode in this application scenario. Figure 10cThis diagram illustrates the application's "swipe activation + swipe" interaction mode within a given scenario. In this application scenario, the target device can be adapted to... Figure 1 The implementation environment is shown. Please refer back to [link / reference]. Figure 1 The target device 130 can be a PTZ camera (hereinafter referred to as "PTZ"). Based on the interaction between the PTZ and the user terminal 110 and the gateway 150, the image processing method is implemented.

[0179] like Figure 10a As shown, firstly, the real-time video feed of the PTZ is displayed on the user terminal, and user gestures touched on the real-time video feed are detected in real time.

[0180] When a long press is detected, the area selection function is activated, such as... Figure 10b The blue circular box shown. Of course, users can also directly click the area selection button to activate the area selection function, such as... Figure 10c The blue circular frame shown. In Figure 10b In version 10c, the blue circular box indicates that the area selection function is active, allowing users to reselect a region of the image on the live video feed.

[0181] If the area selection function is not activated, it is assumed that the user has other intentions and the area location process will not be executed.

[0182] After the area selection function is activated by long-pressing or directly clicking the area selection button, the area positioning process begins. Specifically, a rectangular selection area is drawn in real-time between the touch start point and the current touch point, and the position of the current touch point is dynamically adjusted as the user's finger moves, thereby adjusting the size of the rectangular selection area accordingly (e.g., ...). Figure 10b (Or the blue rectangle shown in 10c). The touch start point and the current touch point serve as a set of diagonal vertices of the rectangular selection area. Each side of the rectangular selection area is parallel to the horizontal and vertical boundaries of the real-time video frame, forming an axis-aligned rectangle.

[0183] When the user releases the touch (the rectangular selection area disappears), that is, when the user raises their hand, the position at the time of hand release is taken as the touch end point. Based on the touch start point and touch end point, the coordinates and size of the final rectangular selection area are obtained. Specifically, based on the center position of the rectangular selection area, the horizontal and vertical angles that the pan-tilt unit needs to rotate are automatically calculated; and based on the size and aspect ratio of the rectangular selection area, the optimal scaling factor is automatically calculated so that the magnified real-time video image can completely surround the rectangular selection area, and the rectangular selection area is located at the center of the real-time video image.

[0184] Finally, the gateway generates device control commands based on the required horizontal and vertical angles of the PTZ (pan-tilt-zoom) rotation and / or zoom level, and sends these commands to the PTZ to synchronously execute rotation and zoom actions (if the limits are exceeded, it will rotate to the limit position or zoom to the limit factor). The images captured by the PTZ after synchronous rotation and zoom actions are then displayed on the user's terminal. Figure 10b As shown in 10c, a new live video feed is displayed on the user terminal. This new live video feed is an enlarged image that completely surrounds the rectangular selection area.

[0185] In addition, Figure 10b In version 10c, the zoom level displayed below the real-time video display area can include two scenarios: (1) Automatic update based on screen adjustment: The zoom level text on the current zoom level button will be refreshed in real time to reflect the current zoom level, meaning the zoom level text will be updated in real time after the user adjusts the zoom level; (2) Update based on user manual adjustment: If the user manually clicks the left button among the three buttons, the screen will become 1x, the left button will become selected (the circle radius is larger in the selected state compared to the unselected state), and the zoom level text will also become 1x. Finally, the zoomed real-time video screen is displayed on the user's terminal.

[0186] At the same time, such as Figure 10b As shown in 10c, when the area selection function is not activated, users can also directly control the gimbal to perform rotation by clicking the up, down, left, and right buttons in the steering wheel control. Finally, the real-time video image after rotation is displayed on the user terminal.

[0187] In this application scenario, by innovatively introducing a region selection interaction method, users can directly select the corresponding rectangular selection area on the real-time video screen. Then, the spatial position information of the rectangular selection area is automatically converted into the rotation angle of the gimbal. At the same time, the optimal scaling factor is intelligently calculated based on the size and shape characteristics of the rectangular selection area, realizing integrated control of "positioning + scaling". This simplifies the traditional 6-10 step operation process to a single step, which not only improves operational efficiency but also reduces the number of network interactions from multiple round trips to one, significantly improving the user experience.

[0188] Compared with existing technologies, the embodiments of this application have the following significant advantages: First, significantly improved operational efficiency. By integrating area selection, coordinate mapping, angle calculation, and zoom factor calculation into a unified control process, the user only needs one interaction from initiating the selection operation to obtaining the target image. The number of operations is reduced from 6 to 10 in the prior art to 1, and the operation time is shortened from 10 to 15 seconds to 2 to 3 seconds, improving efficiency by more than 5 times. Second, intelligent calculation of the optimal zoom factor. The system automatically selects to calculate the zoom factor based on the relationship between the aspect ratio of the selected area and the aspect ratio of the image, ensuring that the magnified image can completely surround the selected area and that the selected area is located in the center of the image, completely solving the user's cognitive burden of "not knowing how many times to zoom in". Third, intuitive and natural operation. The selection operation conforms to the user's habit of using their finger to select the target area in daily life, with extremely low learning cost. Users do not need to understand the technical principles of gimbal rotation and digital zoom to complete the operation. Fourth, solving the problem of accumulated network latency. By encapsulating the pan-tilt rotation angle and zoom level into a single control command and sending it synchronously to the pan-tilt camera, the latency caused by multiple network round trips due to step-by-step transmission is avoided, significantly improving the user experience in environments with poor network conditions; fifth, it has a wide range of applications. This application is applicable to all application scenarios requiring remote monitoring and detailed magnification, such as home surveillance, commercial security, infant monitoring, smart homes, industrial inspection, and education and training, demonstrating good versatility.

[0189] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0190] The following are embodiments of the apparatus described in this application, which can be used to execute the image processing method involved in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the method embodiments of the image processing method involved in this application.

[0191] Please see Figure 11 This application provides an image processing device 900, including but not limited to: a first display module 910, a first data acquisition module 930, and a second display module 950.

[0192] The first display module 910 is used to display a first image; the first image is an image captured and collected by the target device in the target space.

[0193] The first data acquisition module 930 is used to acquire first position data if an operation triggered in the first screen to determine a screen area is detected; the first position data is used to indicate the position of the screen area in the first screen.

[0194] The second display module 950 is used to display a second image obtained based on the first position data; the second image is an image obtained by the target device after adjusting the camera angle and / or focal length based on the second position data, and capturing and collecting images of a spatial area; the second position data is obtained by mapping the first position data and is used to indicate the position of the spatial area in the target space under the field of view of the target device.

[0195] In an exemplary embodiment, the first data acquisition module 930 is further configured to determine whether a first gesture operation is detected on the first screen based on the point of action on the first screen; the first gesture operation is used to indicate the selection of a screen area in the first screen; if so, the screen area is determined in the first screen based on the start and end positions of the point of action of the first gesture operation on the first screen; and the first position data is obtained based on the position of the screen area in the first screen.

[0196] In an exemplary embodiment, the first data acquisition module 930 is further configured to track the movement trajectory of the action point on the first screen when an action point is detected on the first screen; if the length of the movement trajectory exceeds a distance threshold, then it is determined that a first gesture operation has been detected.

[0197] In an exemplary embodiment, the first data acquisition module 930 is further configured to determine whether the region selection function is enabled; the region selection function is used to indicate whether it is allowed to select a region in the first screen; if yes, then the step of acquiring first location data is performed if an operation triggered to determine a region in the first screen is detected.

[0198] In an exemplary embodiment, the device further includes a function determination module, used to detect on the first screen whether the dwell time of the point of action at the current position exceeds a time threshold; if so, the mark area selection function has been enabled.

[0199] In an exemplary embodiment, the function determination module is further configured to mark the region determination control as active when the region selection function is enabled; the active region determination control is used to indicate that a region of the screen can be selected in the first screen.

[0200] In an exemplary embodiment, the function determination module is further configured to determine whether an operation triggered by switching the region determination control from an inactive state to an active state is detected; if so, the region determination control is marked as active to indicate that the selection of a screen region in the first screen is allowed.

[0201] In an exemplary embodiment, the second display module 950 is further configured to, if no operation triggered for determining the screen area is detected on the first screen, instruct the target device to position the focus at the location based on the position of the point of action when it stays on the first screen; and display a third screen; the third screen is a screen captured and collected by the target device after positioning in the target space.

[0202] In an exemplary embodiment, the device further includes a screen adjustment module, configured to acquire device adjustment data if an operation triggered by adjusting the camera angle and / or focal length of the target device is detected; the device adjustment data is used to indicate the camera angle and / or focal length to be adjusted of the target device; and to display a fourth screen acquired based on the device adjustment data; the fourth screen is a screen captured and collected by the target device after adjusting the camera angle and / or focal length based on the device adjustment data, targeting the target space.

[0203] In an exemplary embodiment, the image adjustment module is further configured to, when the region determination control is in an active state, mark the region determination control as inactive when an operation triggered on the region determination control is detected, so as to detect whether there is an operation triggered by adjusting the camera angle and / or focal length of the target device when the region determination control is inactive.

[0204] Please see Figure 12 This application provides a device control device 1000, including but not limited to: a second data acquisition module 1010, a data mapping module 1030, and a device control module 1050.

[0205] The second data acquisition module 1010 is used to receive first position data; the first position data is used to indicate the position of the screen area in the first screen; the screen area is determined in response to the operation triggered in the first screen to determine the screen area; the first screen is the screen captured and collected by the target device for the target space.

[0206] The data mapping module 1030 is used to map the first position data from the coordinate system where the first screen is located to the coordinate system where the target device is located to obtain the second position data; the second position data is used to indicate the position of the spatial region in the target space under the field of view of the target device.

[0207] The device control module 1050 is used to control the target device to adjust the camera angle and / or focal length based on the second position data, so that the target device can take pictures and collect a second image of the spatial area after adjusting the camera angle and / or focal length.

[0208] In an exemplary embodiment, the second data acquisition module 1010 is further configured to calculate the camera angle and / or focal length to be adjusted of the target device based on the first position data; and obtain the second position data from the camera angle and / or focal length to be adjusted of the target device.

[0209] In an exemplary embodiment, the second data acquisition module 1010 is further configured to determine the center position of the screen area based on the first position data; calculate the deviation between the center position of the screen area and the center position of the first screen; and calculate the camera angle to be adjusted of the target device based on the deviation and the field of view of the target device.

[0210] In an exemplary embodiment, the second data acquisition module 1010 is further configured to determine the width, height, and aspect ratio of the image area based on the first position data; the aspect ratio refers to the ratio of the width to the height of the image area; if the aspect ratio is less than the image aspect ratio, the focal length to be adjusted for the target device is calculated based on the height of the first image and the height of the image area; the image aspect ratio refers to the ratio of the width to the height of the first image; if the aspect ratio is greater than or equal to the image aspect ratio, the focal length to be adjusted for the target device is calculated based on the width of the first image and the width of the image area.

[0211] In an exemplary embodiment, the above-described apparatus further includes a prompting module, configured to, if the camera angle and / or focal length to be adjusted of the target device exceeds the adjustment range of the target device, restrict the camera angle and / or focal length to be adjusted of the target device based on the adjustment range of the target device, so that the camera angle and / or focal length to be adjusted of the target device does not exceed the adjustment range of the target device; and send a second prompting message; the second prompting message is used to prompt that the adjusted camera angle and / or focal length of the target device has reached the adjustment range.

[0212] It should be noted that the image processing device provided in the above embodiments is only illustrated by the division of the above functional modules when performing image processing. In actual applications, the above functions can be assigned to different functional modules as needed. That is, the internal structure of the image processing device will be divided into different functional modules to complete all or part of the functions described above.

[0213] Furthermore, the image processing apparatus and image processing method embodiments provided in the above embodiments belong to the same concept, and the specific way in which each module performs operations has been described in detail in the method embodiments, and will not be repeated here.

[0214] Please see Figure 13 This application provides an electronic device 4000, which may include: smartphones, tablets, laptops, desktop computers, smart control panel gateways, servers, etc.

[0215] exist Figure 13 In this context, the electronic device 4000 includes at least one processor 4001 and at least one memory 4003.

[0216] Data interaction between the processor 4001 and the memory 4003 can be achieved through at least one communication bus 4002. This communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used to represent it in the figure, but this does not indicate that there is only one bus or one type of bus.

[0217] Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.

[0218] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0219] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device capable of storing static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device capable of storing information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing computer programs having instruction or data structure forms and accessible by the electronic device 4000, but not limited to these.

[0220] The memory 4003 stores a computer program, and the processor 4001 can read the computer program stored in the memory 4003 through the communication bus 4002.

[0221] The computer program is executed by one or more processors 4001 to implement the methods in the above embodiments.

[0222] Furthermore, this application provides a storage medium storing a computer program that is executed by one or more processors to implement the method described above.

[0223] This application provides a computer program product, including a computer program that is executed by one or more processors to implement the method described above.

[0224] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. An image processing method, characterized in that, The method includes: The first screen is displayed; the first screen is the image captured and collected by the target device in the target space; If an operation triggered to define a screen area is detected in the first screen, first location data is acquired; the first location data is used to indicate the position of the screen area in the first screen. The second image is displayed based on the first location data; the second image is the image captured and collected by the target device after adjusting the camera angle and / or focal length based on the second location data for a spatial region; the second location data is mapped from the first location data and is used to indicate the position of the spatial region in the target space under the field of view of the target device.

2. The method as described in claim 1, characterized in that, If an operation triggered in the first screen for a defined screen area is detected, obtaining the first location data includes: Based on the point of action on the first screen, it is determined whether a first gesture operation is detected on the first screen; the first gesture operation is used to indicate the selection of the screen area in the first screen; If so, the screen area is determined in the first screen based on the start and end positions of the point of action of the first gesture on the first screen; The first position data is obtained based on the position of the image area in the first image.

3. The method as described in claim 2, characterized in that, The step of determining whether a first gesture operation is detected on the first screen based on the point of action on the first screen includes: When the point of action is detected on the first screen, the movement trajectory of the point of action on the first screen is tracked; If the length of the movement trajectory exceeds a distance threshold, then the first gesture operation is detected.

4. The method as described in claim 1, characterized in that, Before acquiring the first location data, if an operation triggered to define a screen area is detected in the first screen, the method includes: Determine whether the region selection function is enabled; the region selection function is used to indicate whether it is allowed to select the region of the screen in the first screen; If so, then the step of obtaining the first location data is performed if an operation triggered in the first screen for a defined screen area is detected.

5. The method as described in claim 4, characterized in that, Whether the function for determining the selected region is enabled includes: On the first screen, the detection checks whether the dwell time of the detection point at the current position exceeds a time threshold; If yes, then the region selection function is marked as enabled.

6. The method as described in claim 5, characterized in that, After determining that the region selection function has been enabled, the method further includes: If the area selection function is confirmed to be enabled, the area selection control is marked as active; the active area selection control is used to indicate that the area of ​​the screen can be selected in the first screen.

7. The method as described in claim 4, characterized in that, Whether the function for determining the selected region is enabled includes: Determine whether an operation triggered by switching the region definition control from an inactive state to an active state was detected; If so, the area determination control is marked as active to indicate that the screen area can be selected in the first screen.

8. The method according to any one of claims 1 to 7, characterized in that, After displaying the first screen, the method further includes: If no operation triggered to determine the screen area is detected on the first screen, the target device is instructed to position the focus to the position based on the position of the point of action when it stays on the first screen. The third screen is displayed; the third screen is the image captured and collected by the target device after positioning in the target space.

9. The method according to any one of claims 1 to 7, characterized in that, After displaying the first screen, the method further includes: If an operation triggered by adjusting the camera angle and / or focal length of the target device is detected, device adjustment data is acquired; the device adjustment data is used to indicate the camera angle and / or focal length of the target device to be adjusted. The fourth screen is displayed based on the adjustment data obtained by the device; the fourth screen is the image captured and collected by the target device after adjusting the camera angle and / or the focal length based on the device adjustment data, targeting the target space.

10. The method as described in claim 9, characterized in that, Before acquiring the second device adjustment data, if an operation triggered by adjusting the camera angle and / or focal length for the target device is detected, the method further includes: When the region determination control is active, if an operation triggered for the region determination control is detected, the region determination control is marked as inactive. In order to detect whether the operation triggered by adjusting the camera angle and / or focal length for the target device exists when the region determination control is inactive.

11. A device control method, characterized in that, The method includes: Receive first location data; the first location data is used to indicate the position of a screen area in a first screen; the screen area is determined in response to an operation triggered in the first screen to determine the screen area; the first screen is a screen captured and collected by the target device in relation to a target space; The first position data is mapped from the coordinate system where the first screen is located to the coordinate system where the target device is located to obtain the second position data; the second position data is used to indicate the position of the spatial region in the target space under the field of view of the target device. Based on the second location data, the target device is controlled to adjust the camera angle and / or focal length so that after adjusting the camera angle and / or focal length, the target device can capture and obtain a second image of the spatial area.

12. The method as described in claim 11, characterized in that, The step of mapping the first location data from the coordinate system of the first screen to the coordinate system of the target device to obtain the second location data includes: Based on the first location data, calculate the camera angle and / or focal length to be adjusted for the target device; The second position data is obtained from the camera angle and / or focal length to be adjusted of the target device.

13. The method as described in claim 12, characterized in that, The step of calculating the camera angle and / or focal length to be adjusted for the target device based on the first location data includes: Based on the first location data, the center position of the image area is determined; Calculate the deviation between the center position of the image area and the center position of the first image; Based on the deviation and the field of view of the target device, the camera angle to be adjusted of the target device is calculated.

14. The method as described in claim 12, characterized in that, The step of calculating the camera angle and / or focal length to be adjusted for the target device based on the first location data includes: Based on the first location data, the width, height, and aspect ratio of the image area are determined; the aspect ratio refers to the ratio of the width to the height of the image area. If the aspect ratio of the region is less than the aspect ratio of the image, then the focal length to be adjusted for the target device is calculated based on the height of the first image and the height of the image region; the aspect ratio refers to the ratio of the width to the height of the first image. If the aspect ratio of the region is greater than or equal to the aspect ratio of the image, then the focal length to be adjusted for the target device is calculated based on the width of the first image and the width of the image region.

15. The method as described in claim 12, characterized in that, Before obtaining the second position data from the camera angle and / or focal length to be adjusted by the target device, the method further includes: If the camera angle and / or focal length to be adjusted of the target device exceeds the adjustment range of the target device, then based on the adjustment range of the target device, the camera angle and / or focal length to be adjusted of the target device are restricted so that the camera angle and / or focal length to be adjusted of the target device does not exceed the adjustment range of the target device; Send a second notification message; the second notification message is used to indicate that the adjusted camera angle and / or focal length of the target device has reached the adjustment range.

16. An electronic device comprising at least one processor and at least one memory, wherein, The memory stores a computer program, characterized in that the computer program, when executed by the processor, implements the method as described in any one of claims 1 to 15.

17. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by one or more processors, it implements the method as described in any one of claims 1 to 15.

18. A computer program product comprising a computer program, characterized in that, When the computer program is executed by one or more processors, it implements the method as described in any one of claims 1 to 15.