Mobile control method of mobile device and mobile device, storage medium

CN122756239APending Publication Date: 2026-09-15BRIGHTWAY INNOVATION INTELLIGENT TECH (SUZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610894198.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-15

Smart Images

  • Figure CN122756239A_ABST
    Figure CN122756239A_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a mobile control method of a mobile device, the mobile device, and a storage medium, wherein the mobile device comprises a multi-vision system, the multi-vision system comprises at least two image acquisition components, the method comprises the following steps: positioning a tracking target, adjusting a field of view range of a target image acquisition component based on historical motion information of the tracking target, and searching the tracking target from an acquired image of any image acquisition component in the multi-vision system, in the case that the tracking target is searched from the acquired image of at least one image acquisition component in the multi-vision system, the mobile device is controlled to move along the relative direction of the tracking target, so that the target recapture is completed with the minimum motion cost on the premise that the mobile device itself is not burdened with the mobility, the time required for the target recapture is significantly shortened, and the tracking failure rate caused by the limited mobility of the mobile device is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of mobile control, and more specifically, to a mobile control method for a mobile device, a mobile device, and a storage medium. Background Technology

[0002] In intelligent mobile devices such as self-following scooters, robust tracking of specific targets in complex scenarios requires pure visual recognition. Mobile devices (such as electric two-wheelers) generally use forward-fixed vision systems with limited field of view. When the tracked target is temporarily obscured by pedestrians or obstacles, or moves to the edge of the vision system's field of view (such as turning right or moving backward, causing the tracked target to move out of the field of view), the fixed-oriented vision system cannot actively adjust its observation direction and can only rely on the overall turning of the mobile device to recapture it. Therefore, in related technologies, there is a technical problem of low efficiency in searching for and tracking targets by mobile devices due to the limitations of field of view and movement flexibility when the tracked target disappears from the vision system's field of view. Summary of the Invention

[0003] This application provides a mobile control method for a mobile device, as well as the mobile device and storage medium, to at least solve the technical problem of low efficiency in searching and tracking targets by mobile devices due to limitations in field of view and movement flexibility in related technologies.

[0004] According to one aspect of the embodiments of this application, a motion control method for a mobile device is provided, executed by the mobile device, the mobile device including a multi-view vision system, the multi-view vision system including at least two image acquisition components; the method includes: continuously locating a tracking target based on a sequence of acquired image groups from the multi-view vision system, wherein each acquired image group in the sequence includes at least two acquired images with matching acquisition times; in response to the tracking target being in a lost state, adjusting the field of view of a target image acquisition component based on historical motion information of the tracking target, and searching for the tracking target from the acquired images of any image acquisition component in the multi-view vision system, wherein the target image acquisition component is the image acquisition component that last identified the tracking target before the tracking target was in the lost state; and, if the tracking target is found from the acquired images of at least one image acquisition component in the multi-view vision system, performing motion control on the mobile device along the relative direction of the tracking target.

[0005] According to another aspect of the embodiments of this application, a mobile device is also provided, comprising: a multi-view vision system and a control unit, the multi-view vision system including at least two image acquisition units; wherein the multi-view vision system is used for image acquisition; the control unit is used for continuously locating a tracking target based on a sequence of acquired image groups from the multi-view vision system, wherein each acquired image group in the sequence includes at least two acquired images with matching acquisition times; in response to the tracking target being in a lost state, adjusting the field of view of the target image acquisition unit based on the historical motion information of the tracking target, and searching for the tracking target from the acquired images of any image acquisition unit in the multi-view vision system, wherein the target image acquisition unit is the image acquisition unit that last identified the tracking target before the tracking target was in the lost state; and when the tracking target is found from the acquired images of at least one image acquisition unit in the multi-view vision system, performing motion control on the mobile device along the relative direction of the tracking target.

[0006] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to perform the steps in any of the above method embodiments when executed by a processor.

[0007] According to another aspect of the embodiments of this application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the steps in any of the method embodiments described above.

[0008] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to perform the steps of any of the above method embodiments through the computer program.

[0009] This application discloses a mobile device comprising a multi-view vision system, which includes at least two image acquisition components. Based on the historical motion information of the tracked target, the field of view of the target image acquisition component is adjusted to directionally expand the search range. The target is searched for from the images acquired by any one of the image acquisition components in the multi-view vision system, achieving efficient local search for the tracked target. When the tracked target is found in the images acquired by at least one of the image acquisition components in the multi-view vision system, the mobile device is moved along the relative direction of the tracked target. This allows for target re-acquisition with minimal motion cost without increasing the mobility burden of the mobile device itself, significantly shortening the time required for target re-acquisition and reducing the tracking failure rate caused by the limited mobility of the mobile device. This improves the efficiency of the mobile device in searching for and tracking targets, thus solving the technical problem of low efficiency in searching for and tracking targets by mobile devices due to limited field of view and insufficient motion flexibility in related technologies. Attached Figure Description

[0010] Figure 1 This is a schematic diagram illustrating an application scenario of a mobile device control method according to an embodiment of this application;

[0011] Figure 2 This is a schematic diagram of the hardware structure of a mobile device control method according to an embodiment of this application;

[0012] Figure 3 This is a flowchart illustrating an optional mobile device movement control method according to an embodiment of this application.

[0013] Figure 4 This is a schematic diagram of an optional control method according to an embodiment of this application;

[0014] Figure 5 This is a structural block diagram of an optional mobile device according to an embodiment of this application. Detailed Implementation

[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices. According to one aspect of an embodiment of this application, a mobile device motion control method is provided. Optionally, in this embodiment, the above-described mobile device motion control method may be applied to, but is not limited to, […]. Figure 1 The hardware environment shown includes a mobile device 102, a control terminal 104, and a server 106. The mobile device 102 may have network connectivity. The server 106 can connect to the mobile device 102 via a network and can be used to provide services (e.g., application services, location services, etc.) to the mobile device 102 or clients installed on the mobile device 102. A database can be set up on or independently of the server 106 to provide data storage services for the server 106.

[0017] The control terminal 104 can be a mobile terminal or controller bound to the mobile device 102. The mobile device 102 can communicate with the control terminal 104 via a wireless network, which can include, but is not limited to, at least one of the following: Wireless Fidelity (WIFI) and Bluetooth. Optionally, the control terminal 104 can also be directly connected to the mobile device 102 via a data cable or other connecting cable to transmit signals and realize interaction between the control terminal 104 and the mobile device 102.

[0018] Alternatively, the aforementioned mobile device motion control method can also be applied to, but is not limited to, mobile devices such as... Figure 2 The hardware environment of the mobile device 200 shown. The mobile device 200 includes a multi-view vision system 202 and a control unit 204, wherein the multi-view vision system 202 includes at least two image acquisition units 2022.

[0019] Mobile devices can be, but are not limited to, electric two-wheeled vehicles. Optionally, electric two-wheeled vehicles can include electric scooters, electric balance scooters, electric bicycles, etc. Electric two-wheeled vehicles can include components such as the frame, front wheel, rear wheel, and seat, and can also include a Vehicle Control Unit (VCU), a complex system integrating hardware and software. For the hardware portion, the VCU can include one or more hardware microprocessors, microcontroller units (MCUs), and necessary input / output interfaces, memory, power modules, communication modules, etc. These hardware components constitute the physical basis of the VCU, enabling it to receive signals, process data, send control commands, and communicate with other vehicle subsystems or external devices. For the software portion, the VCU's software can include an embedded operating system, application programs, control algorithms, and diagnostic programs. The software is responsible for parsing data from sensors and subsystems, performing complex calculations and logical judgments, and generating control signals for actuators. The software is typically written as a program capable of handling functions such as vehicle powertrain, energy management, and safety control.

[0020] The mobile device movement control method of this application embodiment can be executed by the mobile device 202. Taking the execution of the mobile device movement control method of this embodiment by the mobile device 202 as an example, Figure 3 This is a flowchart illustrating an optional mobile device movement control method according to an embodiment of this application, as shown below. Figure 3 As shown, the process of this method may include the following steps S302 to S306.

[0021] Step S302: Based on the sequence of acquired images from the multi-view vision system, the target is continuously located. Each acquired image group in the sequence includes at least two acquired images with matching acquisition times.

[0022] Step S304: In response to the tracking target being lost, based on the historical motion information of the tracking target, the field of view of the target image acquisition component is adjusted by the adjustment mechanism, and the tracking target is searched from the acquired images of any image acquisition component in the multi-view vision system. The target image acquisition component is the image acquisition component that last identified the tracking target before the tracking target was lost.

[0023] Step S306: If a tracking target is found in the image acquired by at least one image acquisition component of the multi-view vision system, the mobile device is moved along the relative direction of the tracking target.

[0024] The mobile device movement control method in this embodiment can be applied to the field of movement control technology and to scenarios involving movement control of mobile devices. Optionally, when the mobile device is an electric two-wheeled vehicle, the electric two-wheeled vehicle can track a target. By recognizing and searching for the target through the vehicle's vision system, the vehicle can be controlled to track the target, thereby achieving automatic tracking of the target by the electric two-wheeled vehicle.

[0025] Mobile devices (such as electric two-wheelers) generally use forward-facing fixed vision systems, which have a limited field of view. Furthermore, forward-facing fixed vision systems suffer from problems such as small overlap areas, large blind spots, and easy target loss. When the tracked target is obscured for a long time or enters a blind spot (such as when turning right or moving backward, causing the tracked target to move out of the field of view), the fixed-facing vision system cannot actively adjust the observation direction and can only rely on the overall turning of the mobile device to re-capture it. However, the mobility of mobile devices (such as electric scooters) is limited, and they cannot quickly rotate to enrich the field of view. Therefore, in related technologies, there is a technical problem of low efficiency in searching for and tracking targets by mobile devices when the tracked target disappears from the field of view of the vision system due to the limitations of field of view and mobility.

[0026] To at least partially solve the aforementioned technical problems, in this embodiment, the mobile device includes a multi-view vision system, which includes at least two image acquisition components. Based on the historical motion information of the tracked target, the field of view of the target image acquisition component is adjusted to directionally expand the search range, thereby achieving efficient local search of the tracked target. When the tracked target is found in the image acquired by at least one image acquisition component of the multi-view vision system, the mobile device is moved along the relative direction of the tracked target. This allows for target re-acquisition with minimal motion cost without increasing the mobility burden of the mobile device itself, significantly shortening the time required for target re-acquisition and reducing the tracking failure rate caused by limited device mobility, thereby improving the efficiency of the mobile device in searching for and tracking targets.

[0027] In this embodiment, the mobile device refers to a mobile device with autonomous mobility. The mobile device includes a multi-view vision system. Optionally, the mobile device can be an electric two-wheeled vehicle, which can include electric scooters, electric balance bikes, and electric bicycles, etc.

[0028] A multi-view vision system refers to a visual perception system that includes two or more image acquisition components. Optionally, a multi-view vision system includes at least two image acquisition components. One of the at least two image acquisition components refers to a physical device that acquires visual information for continuous localization of the tracked target. Optionally, the image acquisition component can be a camera. The tracked target refers to the object that the mobile device follows. Optionally, the tracked target can be the user of the mobile device, a target, a controller, or a lane, etc.

[0029] It should be noted that in this embodiment, each image acquisition component in the multi-view vision system can be controlled independently, meaning that each image acquisition component does not interfere with the others.

[0030] In some embodiments, a multi-view vision system can acquire multiple images to obtain a sequence of acquired images. The sequence of acquired images refers to a set of acquired image groups acquired sequentially in chronological order. Optionally, the sequence includes multiple acquired image groups, and each acquired image group may contain at least two time-synchronized (i.e., time-matched) acquired images from different image acquisition components of the multi-view vision system; that is, each acquired image group includes at least two time-matched acquired images.

[0031] Each image group in the image acquisition sequence refers to a group of images synchronously acquired by at least two image acquisition components in a multi-view vision system at the same acquisition time. Each image group in the image acquisition sequence includes at least two acquired images with matching acquisition times, and the at least two acquired images with matching acquisition times correspond one-to-one with at least two image acquisition components.

[0032] In some embodiments, the tracking target in each group of acquired images can be detected and tracked in real time, and the position information of the tracking target in the image coordinate system of the multi-view vision system can be continuously output through feature matching or target recognition algorithms (such as YOLO, SORT, etc.). Optionally, the sequence of acquired images acquired by the multi-view vision system is obtained, and the tracking target in the two acquired images in each group of acquired images in the sequence of acquired images is detected to obtain the first detection box coordinates and the second detection box coordinates of the tracking target; according to the focal length of the image acquisition component, the horizontal physical distance between the two image acquisition components and the disparity, the three-dimensional distance from the following target to the mobile device is calculated, so as to obtain the position of the following target relative to the tracking target. The focal length of each image acquisition component is the same, and the disparity refers to the difference between the first detection box coordinates and the second detection box coordinates of the tracking target point in the two acquired images. For example, the three-dimensional distance from the following target to the mobile device can be calculated by the following formula (1):

[0033] D= (1)

[0034] Where D is the three-dimensional distance from the target to the mobile device, that is, D can be the straight-line distance from the target along the optical axis to the midpoint of the horizontal physical distance between the two image acquisition components; f is the focal length of the image acquisition component; B is the horizontal physical distance between the two image acquisition components; and d is the parallax.

[0035] A lost target state refers to a state where the tracked target is not detected in any of the currently acquired images from the multi-view vision system. Optionally, when the tracked target is in a lost target state, it indicates that the tracked target has left the field of view of each image acquisition component in the multi-view vision system. Optionally, the tracked target is determined to be in a lost target state when no tracked target is detected in any of the acquired images in the current image sequence. In the case of a lost target state, the multi-view vision system suspends binocular ranging and resumes target tracking.

[0036] Historical motion information of the tracked target refers to the motion characteristics of the target in multiple consecutive frames before it is lost. Optionally, historical motion information may include position information, direction of motion, speed, distance between the tracked target and the edge of the acquired image, and parallax change rate (i.e., the trend of depth change of the tracked target).

[0037] The target image acquisition component is the last image acquisition component to identify the tracked target before the tracked target is lost.

[0038] Any image acquisition component in a multi-view vision system refers to any image acquisition component in the multi-view vision system. Optionally, any image acquisition component in a multi-view vision system can be a target image acquisition component or other image acquisition components besides the target image acquisition component.

[0039] Optionally, if one image acquisition component in the multi-view vision system identifies the target before the target is lost, while other image acquisition components do not identify the target, then the image acquisition component that identifies the target is the target image acquisition component.

[0040] In some embodiments, after acquiring the historical motion information of the tracked target, the field of view of the target image acquisition component can be adjusted by a controller; alternatively, an adjustment mechanism can be set on the mobile device to adjust the field of view of the target image acquisition component. The adjustment mechanism refers to a mechanical mechanism used to physically change the spatial pose of the image acquisition component, and can be used to adjust the field of view of each image acquisition component in the multi-view vision system separately. Optionally, the adjustment mechanism may include a translation rail, a slider mechanism, a rotary motor, and a drive controller, etc.; the field of view of the target image acquisition component itself can also be adjusted. In some embodiments, before the tracked target is lost, two image acquisition components in the multi-view vision system both identify the tracked target. Based on the historical motion information of the tracked target, the image acquisition component matching the historical motion information is selected as the target image acquisition component. For example, assuming that the tracked target is located on the right side of the acquired image, the target image acquisition component can be the image acquisition component installed on the right side of the multi-view vision system. As another example, assuming that the movement direction of the tracked target is to the right, the target image acquisition component can be the image acquisition component installed on the right side of the multi-view vision system.

[0041] Optionally, the target image acquisition component can be an image acquisition device that matches the position information of the tracked target, an image acquisition device that matches the movement direction of the tracked target, or an image acquisition device that matches both the position information and the movement direction of the tracked target.

[0042] In some embodiments, assuming the multi-view vision system includes a first image acquisition unit and a second image acquisition unit, before the tracked target is lost, both the images acquired by the first image acquisition unit and the images acquired by the second image acquisition unit identify the tracked target. Then, the tracking confidence of the first image acquisition unit and the tracking confidence of the second image acquisition unit are obtained, as well as the disparity variance of the first image acquisition unit and the disparity variance of the second image acquisition unit. If the tracking confidence of the first image acquisition unit is greater than the tracking confidence of the second image acquisition unit and the disparity variance of the first image acquisition unit is less than the disparity variance of the second image acquisition unit, it is determined that the state of the first image acquisition unit is more stable than the state of the second image acquisition unit. The first image acquisition unit is matched with historical motion information, and the first image acquisition unit is determined to be the target image acquisition unit.

[0043] The field of view of the target image acquisition component refers to the range of physical spatial angles that the target image acquisition component can cover. Optionally, the field of view of the target image acquisition component can be determined based on the optical lens parameters and orientation of the target image acquisition component, and can be adjusted by translation or rotation. For example, adjusting the field of view of the target image acquisition component can be achieved by adjusting the horizontal position of the target image acquisition component, thereby adjusting the field of view; or by rotating the target image acquisition component using the adjustment mechanism, thereby adjusting the field of view.

[0044] In some embodiments, after determining that the tracked target is lost, historical motion information of the tracked target is acquired, candidate target regions (i.e., the possible locations where the tracked target may appear) are determined based on the historical motion information, and the field of view of the target image acquisition component that matches the historical motion information is adjusted by the control adjustment mechanism based on the candidate target regions. Each image acquisition component in the multi-view vision system continuously acquires images and searches for the tracked target in the acquired images acquired by each image acquisition component in the multi-view vision system.

[0045] Optionally, the motion direction and position coordinates of the tracked target in the last three frames of acquired images before it was lost are obtained to determine the motion trend of the tracked target. Based on the motion trend, candidate target regions are determined. By adjusting the position of the target image acquisition component that matches the motion direction and position coordinates, the field of view of the target image acquisition component can search for the tracking target. At the same time, other image acquisition components except the target image acquisition component keep their original positions and orientations unchanged and continue to acquire images within the corresponding field of view. In the acquired images acquired by all image acquisition components, the candidate target regions are re-detected and feature-matched (e.g., REID) to identify whether the tracked target has reappeared.

[0046] At least one image acquisition component in a multi-view vision system refers to an image acquisition component that has re-searched for the tracked target. Optionally, at least one image acquisition component can be a target acquisition component, or it can be a target image acquisition component and other image acquisition components besides the target image acquisition component, or it can be at least one image acquisition component besides the target image acquisition component.

[0047] For example, suppose that after adjusting the field of view of the target image acquisition component, the target image acquisition component finds the target to be tracked. In this scenario, at least one image acquisition component includes the target image acquisition component. Then, the mobile device is moved along the relative direction of the target to be tracked. During the movement of the mobile device, the field of view of other image acquisition components in the multi-view vision system remains unchanged or is adjusted. The other image acquisition components continue to search for the target to be tracked within their field of view while following the movement of the mobile device. After the other image acquisition components also find the target to be tracked, the field of view of the target image acquisition component can be adjusted back to the original field of view, or the field of view of the target image acquisition component can be kept unchanged.

[0048] For example, after adjusting the field of view of the target image acquisition component, the mobile device can be moved along the relative direction of the tracking target. Alternatively, the field of view of the image acquisition components other than the target image acquisition component can be adjusted. If the image acquisition components other than the target image acquisition component also find the tracking target, then at least one image acquisition device is the image acquisition component other than the target image acquisition component that finds the tracking target.

[0049] Optionally, if a tracking target is found in the image acquired by at least one image acquisition component in the multi-view vision system, the relative position information (such as azimuth and distance) of the tracking target relative to the mobile device is calculated based on the position of the tracking target in the acquired image (such as the image center offset); based on the relative position information, a control command is generated to drive the motion actuator (such as the motor-driven wheel set) of the mobile device, so that the mobile device is displaced along the relative direction of the tracking target (such as the right front, left rear, etc.).

[0050] In an optional embodiment, taking a multi-view vision system including two image acquisition units as an example, in the case of target loss, the motion trend of the tracking target is calculated based on the displacement vector of the center point of the detection box of the tracking target in the last three acquired image groups, and the most likely direction of movement of the tracking target is determined based on the image boundary information. For example, if the tracking target last appears on the right edge of the screen in an acquired image and its movement direction is downward to the right, it is determined that the tracking target is shifting to the right front. The image acquisition unit corresponding to the acquired image is moved by the adjustment mechanism, while the other image acquisition units remain in place and continue to acquire images; Re-Identification (RE-ID) feature comparison is performed on all detection boxes in the current frame image acquired by each image acquisition unit, and cosine similarity is used to match historical target features. Once a matching result with a similarity exceeding a preset threshold appears in the acquired current frame image, the tracking target re-identification is confirmed to be successful, and a motion control command is immediately sent to the motion control system of the mobile device to make the mobile device move towards the target.

[0051] The embodiments provided in this application provide a mobile device that includes a multi-view vision system. The multi-view vision system includes at least two image acquisition components. Based on the historical motion information of the tracked target, the field of view of the target image acquisition component is adjusted to directionally expand the search range. The target is searched from the images acquired by any one of the image acquisition components in the multi-view vision system, achieving efficient local search for the tracked target. When the tracked target is found in the images acquired by at least one image acquisition component in the multi-view vision system, the mobile device is moved along the relative direction of the tracked target. This allows for target re-acquisition with minimal motion cost without increasing the mobile device's own mobility burden, significantly shortening the time required for target re-acquisition and reducing the tracking failure rate caused by the limited mobility of the mobile device. Therefore, the efficiency of the mobile device in searching for and tracking targets can be improved. Thus, this addresses the technical problem in related technologies where the efficiency of mobile devices in searching for and tracking targets is low due to limited field of view and insufficient motion flexibility.

[0052] In an exemplary embodiment, before the tracked target is lost, the tracked target is visible within the field of view of the target image acquisition component; historical motion information includes motion trend information, which characterizes the target motion trend of the tracked target, and the motion trend information is determined based on the historical coordinates of the tracked target identified from the historical images acquired by the target image acquisition component; adjusting the field of view of the target image acquisition component based on the historical motion information of the tracked target includes: adjusting the field of view of the target image acquisition component according to the target motion trend so that the change in the field of view of the target image acquisition component is consistent with the target motion trend.

[0053] In this embodiment, before the tracking target is lost, the tracking target is visible within the field of view of the target image acquisition component. It can be understood that before the tracking target is lost, the tracking target is within the field of view of the target image acquisition component, and the motion trajectory of the tracking target can be captured by the target image acquisition component.

[0054] Motion trend information refers to information used to characterize the motion trend of a tracked target. Optionally, motion trend information can be vector information that characterizes the change law of the tracking target's motion direction and velocity in the acquired image plane, determined from the historical coordinates of the tracked target in multiple consecutive historical images acquired by the target image acquisition unit before the tracked target was lost. For example, motion trend information may include the displacement vector of the center point of the detection box of the tracked target, the motion direction angle of the tracked target, the motion acceleration of the tracked target, and the edge contact state of the tracked target in the image.

[0055] Historical acquisition images refer to a series of consecutive frames of images captured by the target image acquisition component before the tracked target is lost, in which the tracked target is visible.

[0056] Target motion trend refers to the predictive behavioral trend of the physical motion direction and position of the tracking target relative to the image coordinate system of the target image acquisition component before the tracking target is about to be lost. Optionally, the target motion trend can be used to indicate the spatial path through which the tracking target is most likely to exit the field of view of the target image acquisition component.

[0057] It should be noted that the target movement trend is determined through movement trend information.

[0058] In some embodiments, the target motion trend (such as horizontal yaw angle or lateral displacement vector) is input to the motion control module of the mobile device. Through a preset direction mapping relationship, the horizontal component of the target motion trend is converted into the target yaw angle of the rotary motor, and the lateral component is converted into the target stroke of the slide rail translation motor, generating an executable drive signal. The motion control module outputs the drive signal to the adjustment mechanism (such as the rotary motor and translation motor) that controls the target image acquisition component. The adjustment mechanism drives the target image acquisition component to yaw around the vertical axis according to the drive signal. The adjustment mechanism synchronously or sequentially drives the target image acquisition component to move along the horizontal slide rail. As the optical axis of the target image acquisition component deflects and displaces with the mechanical structure, the coverage area of ​​the optical imaging cone of the target image acquisition component in physical space is translated and tilted as a whole, changing the original field of view of the target image acquisition component, so that the tracking target that disappeared in the field of view of the target image acquisition component falls back into the field of view of the target image acquisition component. In addition, the actual yaw angle and translational displacement of the target image acquisition component can be read in real time by the built-in encoder or attitude sensor and compared with the direction angle of the target motion trend. When the angle between the actual field of view center axis of the target image acquisition component and the target motion vector in the target motion trend is less than the preset tolerance threshold (such as ±5°), it is determined that the change in the field of view range of the target image acquisition component is consistent with the target motion trend.

[0059] In this embodiment, motion trend information is obtained based on the historical coordinates of the tracked target, and the field of view of the target image acquisition component is adjusted according to the target motion trend, so that the change of the field of view of the target image acquisition component is consistent with the target motion trend of the tracked target. Thus, when the tracked target is lost, the tracked target can be searched more quickly.

[0060] In an exemplary embodiment, the motion trend information includes a sequence of movement directions, where a movement direction in the sequence is represented by the coordinate difference between two historical coordinates of the tracked target identified from two adjacent historical acquisition images of the target image acquisition component. Before adjusting the field of view of the target image acquisition component based on the historical motion information of the tracked target, the method further includes: acquiring the latest N movement directions and the latest historical coordinates in the sequence of movement directions, where N is a positive integer greater than or equal to 2, and the latest historical coordinates are the historical coordinates of the tracked target last identified from the historical acquisition images of the target image acquisition component before the current moment; and performing a consistency check on the N movement directions if the distance between the latest historical coordinates and the target image edge in the historical acquisition images of the target image acquisition component is less than or equal to a preset distance threshold, wherein the consistency check is used to check the consistency between the image edge pointed to by the N movement directions and the target image edge, and adjusting the field of view of the target image acquisition component is performed if the consistency check passes.

[0061] In this embodiment, the movement direction sequence refers to a vector set composed of the differences in the historical coordinates of the tracked target in consecutive historical images, arranged in chronological order. Optionally, the movement direction sequence can be used to characterize the tendency of the tracked target's motion path within the field of view of the target image acquisition component. The movement direction sequence can include multiple movement directions, and one movement direction in the sequence can be represented by a displacement vector. Each movement direction in the sequence represents the instantaneous displacement of the tracked target between two adjacent historical images. Here, a movement direction in the movement direction sequence is represented by the coordinate difference between two historical coordinates of the tracked target identified from two adjacent historical images of the target image acquisition component.

[0062] For example, in the historical image acquired in frame T-1, the center pixel coordinates of the detection box of the tracked target are obtained as (X... T-1 Y T-1 = (200, 150); In the historical image of the Tth frame adjacent to this frame, the center coordinates of the same tracked target are updated to (X) = (200, 150); T Y T )=(218, 147); Perform difference calculation on the coordinates of two adjacent frames: ΔX1= X T - X T-1 =18, ΔY1= Y T - Y T-1 =-3, that is, one of the motion directions in the motion direction sequence is (ΔX1, ΔY1) = (18, -3), and (ΔX1, ΔY1) constitute the instantaneous displacement vector in two adjacent historical acquisition images.

[0063] The latest N directions of motion refer to the vectors of the last N consecutive directions of motion in the sequence of movement directions, arranged in chronological order by timestamp.

[0064] The latest historical coordinates refer to the historical coordinates of the tracked target last identified from the historical images acquired by the target image acquisition unit before the current moment. The current moment refers to the moment when the target was identified as being lost.

[0065] The target image edge refers to the geometric boundary corresponding to the effective imaging area in the historically acquired image. Optionally, the target image edge may include the left, right, top, and bottom pixel boundary lines of the historically acquired image. For example, the target image edge may be the edge of the image frame in the historically acquired image.

[0066] The preset distance threshold refers to a preset distance threshold between a latest historical coordinate and the edge of the target image in the historical images acquired by the target image acquisition component. Optionally, if the distance between the latest historical coordinate and the edge of the target image in the historical images acquired by the target image acquisition component is less than or equal to the preset distance threshold, it can be indicated that the tracked target has approached the edge of the target image in the historical images.

[0067] Consistency verification is a logical judgment process based on the spatial pointing relationship of vectors in motion direction. It is used to verify the consistency between the image edges pointed to by N motion directions and the target image edges.

[0068] Optionally, if the image edge pointed to by each of the latest N motion directions is the same as the target image edge, the consistency check is considered passed. For example, if the target image edge is on the right and all N motion directions point to the right, the consistency check is considered passed. Alternatively, if the image edge pointed to by at least M motion directions is the same as the target image edge, the consistency check is considered passed, where M is less than N and M is a positive integer greater than or equal to 2. Only when the consistency check passes will the operation of adjusting the field of view of the target image acquisition component be performed. If the consistency check fails, the operation of adjusting the field of view of the target image acquisition component will not be performed, and the original mechanical pose of the target image acquisition component will be maintained.

[0069] Optionally, a velocity function can be defined to continuously calculate the coordinate transformation of the center point of the detection box of the tracking target in the current frame of the historical image compared to the previous frame of the historical image. If the horizontal and vertical coordinates of the center point of the detection box of the tracking target are both less than the distance from the edge of the screen in the last frame of the historical image, and the movement direction obtained from several consecutive frames of historical images matches the edge position where the detection box of the tracking target was located before it disappeared, then a consistency check is performed. After the check passes, the target image acquisition component (such as a camera) is rotated in the corresponding direction to search for and track the target.

[0070] In some embodiments, the historical image sequence output by the target image acquisition component is continuously read, and the pixel coordinates of the tracked target in two adjacent historical images are extracted. By calculating the horizontal and vertical difference vectors of the pixel coordinates of the tracked target in two adjacent historical images, a sequence of movement directions arranged in chronological order is generated to represent the instantaneous motion trajectory of the tracked target in the historical image plane. The latest N movement direction vectors are extracted from the movement direction sequence, and the position of the tracked target successfully identified in the last historical image before the current acquisition frame is obtained as the latest historical coordinates. The latest historical coordinates are geometrically compared with the edge of the target image acquisition component, and the pixel distance from the latest historical coordinates to the edge of the image (such as the right edge or the bottom edge) is calculated. When the pixel distance from the latest historical coordinates to the edge of the image is less than or equal to a preset distance threshold, it is determined that the tracked target is in a critical state about to leave the field of view of the target image acquisition component. A consistency study is conducted on N consecutive motion directions to verify whether the image boundaries projected by each motion direction are completely consistent. If the image boundaries projected by multiple motion directions are different, the consistency is deemed to have failed. If the image boundaries projected by each motion direction are completely consistent, the consistency check is passed. When the consistency check result is passed, the field of view of the target image acquisition component is adjusted towards the direction in which the tracked target disappears through the adjustment mechanism, and the subsequent target tracking search or re-identification process is initiated.

[0071] In this embodiment, a sequence of movement directions is constructed, and before adjustment is performed, the latest historical coordinates and their corresponding N movement directions near the edge of the target image are checked for consistency. The field of view of the target image acquisition component is adjusted only when the movement trend is confirmed to be stable and pointing to the same edge of the field of view. This can effectively filter out misjudgments caused by instantaneous jitter, occlusion artifacts or random displacement of the tracked target, and significantly improve the accuracy of re-identifying the tracked target after adjusting the field of view of the target image acquisition component.

[0072] In one exemplary embodiment, adjusting the field of view of the target image acquisition component based on historical motion information of the tracked target includes at least one of the following: controlling the target image acquisition component to rotate based on historical motion information to adjust the field of view direction of the target image acquisition component; moving the target image acquisition component along a target slide rail based on historical motion information to adjust the field of view of the target image acquisition component, wherein each image acquisition component is respectively disposed on a slide rail; wherein, in a multi-view vision system, other image acquisition components besides the target image acquisition component remain stationary relative to the mobile device.

[0073] In this embodiment, the field of view direction of the target image acquisition component refers to the spatial azimuth angle pointed to by the optical axis of the target image acquisition component.

[0074] Optionally, adjusting the field of view of the target image acquisition component can be achieved by controlling the target image acquisition component to rotate through an adjustment mechanism, or by controlling the target image acquisition component to move along the target slide rail through an adjustment mechanism, or by controlling the target image acquisition component to rotate or slide through a controller.

[0075] The slide rail refers to a linear guide component fixed on the mobile device. Optionally, the slide rail can provide more accurate translational motion trajectory constraints and mechanical limits for the image acquisition component.

[0076] The target slide rail refers to the slide rail where the target image acquisition component is located. Each image acquisition component is set on a separate slide rail, meaning each image acquisition component has its own corresponding slide rail, thus enabling the image acquisition component to slide. Optionally, the slide rails containing different image acquisition components can be connected together or separated.

[0077] Optionally, the target image acquisition component can be driven to yaw and rotate around its vertical axis by an adjustment mechanism, thereby changing its optical axis orientation and achieving dynamic adjustment of the field of view direction. Alternatively, the target image acquisition component can be driven to slide linearly along its target slide rail by an adjustment mechanism, thereby changing its spatial position and achieving dynamic adjustment of the field of view range (such as baseline length or overlapping area).

[0078] In some embodiments, based on the historical motion information of the tracked target (such as movement direction sequence, latest historical coordinates, and edge distance determination results), the field-of-view compensation requirement of the target image acquisition component is obtained, and corresponding mechanical control commands are generated. When the mechanical control command instructs the target image acquisition component to rotate and adjust, the adjustment mechanism, upon receiving the mechanical control command, drives the target image acquisition component to yaw and rotate around the vertical axis. By changing the physical direction of the optical axis of the target image acquisition component, the field-of-view direction is shifted, thereby bringing the tracked target, which was originally located at the edge of the screen or had moved out of the visible area of ​​the target image acquisition component, back into the field of view of the target image acquisition component. When the mechanical control command instructs the target image acquisition component to translate and adjust, the adjustment mechanism, upon receiving the mechanical control command, drives the target image acquisition component to slide linearly along its target slide rail. By changing the position of the target image acquisition component on the mobile device, the spatial observation point of the target image acquisition component is directly changed, thereby bringing the tracked target, which was originally located at the edge of the screen or had moved out of the visible area of ​​the target image acquisition component, back into the field of view of the target image acquisition component.

[0079] Optionally, during the rotation or movement of the target image acquisition component, an adjustment mechanism controls other image acquisition components (excluding the target image acquisition component) to remain stationary relative to the mobile device. This allows for independent control of the image acquisition components (such as cameras), which can be coordinated with vehicle movement. Alternatively, during the rotation or movement of the target image acquisition component, a controller controls other image acquisition components (excluding the target image acquisition component) to remain stationary relative to the mobile device.

[0080] It should be noted that the target image acquisition component and other image acquisition components are controlled independently. That is, the movement and rotation of the target image acquisition component will not trigger the movement and rotation of other image acquisition components. The movement of the target image acquisition component is driven entirely by control commands based on historical motion information and is not constrained by the motion state of other image acquisition components.

[0081] Optionally, the lateral coordinate sequence and corresponding timestamp of the center point of the detection box of the tracking target in the N consecutive historical images before the tracking target is lost are obtained. The instantaneous lateral velocity vector of the tracking target is calculated by fitting the coordinate sequence and the corresponding timestamp, and the relative positional relationship between the center point of the detection box and the edge of the screen at the moment the tracking target is lost is determined. If the instantaneous lateral velocity vector shows that the tracking target is shifting to the right and finally appears in the right edge area of ​​the screen, the target image acquisition component (e.g., the right camera) is controlled by the adjustment mechanism to rotate the target image acquisition component clockwise around the vertical axis by a preset yaw angle to adjust the field of view of the target image acquisition component, so that the main optical axis of the target image acquisition component is deflected from the front to the prediction search sector on the right side of the original field of view.

[0082] Optionally, the distance between the tracking target and the target image acquisition component is obtained in the historical images of N consecutive frames before the tracking target is lost. If the distance between the tracking target and the target image acquisition component is greater than a first preset distance, the target image acquisition component is controlled to move along the target slide rail by the adjustment mechanism to adjust the field of view of the target image acquisition component.

[0083] In some embodiments, when the tracked target is lost, the movement trend and current possible location of the target are determined by first analyzing the position and direction of movement of the target's detection bounding box in the last few historical images before the target disappeared. For example, if the movement direction in the last few historical images was to the right, and the target disappeared on the right side of the screen in the last historical image, then the right camera is rotated to the right while the left camera remains stationary. Binocular ranging is paused, and each camera performs REID feature matching on all detection bounding boxes in its respective frame to search for the target.

[0084] In this embodiment, based on historical motion information, the target image acquisition component is controlled to rotate or move, while the other image acquisition components remain fixed relative to the mobile device. This allows the field of view adjustment action to more accurately match the lost direction of the tracked target, improving the dynamic adaptability of the tracked target's perspective and the determinism of its adjustment response in occluded or lost scenarios.

[0085] In an exemplary embodiment, adjusting the field of view of the target image acquisition component is achieved by adjusting the field of view direction of the target image acquisition component; during the process of moving the mobile device along the relative direction of the tracking target, the method further includes: continuously identifying the tracking target from the acquired image of any image acquisition component to obtain the target coordinates of the tracking target in the acquired image of any image acquisition component; and, when a preset distance condition is met, stopping the movement control of the mobile device along the relative direction of the tracking target and adjusting the field of view direction of the target image acquisition component back to the specified field of view direction.

[0086] In this embodiment, the target coordinates of the tracking target in the acquired image of any image acquisition component refer to the coordinates of the tracking target in the acquired image of any image acquisition component. Optionally, the target coordinates of the tracking target in the acquired image of any image acquisition component may be the coordinates of the tracking target in the acquired image of the target image acquisition component, or the coordinates of the tracking target in the secondary and tertiary images of other image acquisition components besides the target image acquisition component.

[0087] The preset distance condition refers to the pre-set spatial distance threshold or range between the mobile device and the tracking target. Optionally, the preset distance condition can be used to determine whether the mobile device has ended its movement relative to the tracking target and to trigger the field of view of the target image acquisition component to return to the specified field of view.

[0088] It should be noted that meeting the preset distance condition means that the tracked target has re-entered the effective binocular overlapping field of view of the multi-view vision system, and the tracked target is within the effective parallax range of the multi-view vision system.

[0089] Optionally, if the target coordinates meet a preset distance condition, the movement control of the mobile device along the relative direction of the tracked target is stopped, and the field of view of the target image acquisition component is adjusted back to the specified field of view. For example, if the distance between the target coordinates and the center point coordinates of the currently acquired image is less than a first distance threshold, the movement control of the mobile device along the relative direction of the tracked target is stopped, and the field of view of the target image acquisition component is adjusted back to the specified field of view.

[0090] The specified field of view direction refers to the orientation of the field of view that is initialized or preset. Optionally, the specified field of view direction can be the forward direction corresponding to the moving direction of the mobile device. The specified field of view direction can be used as a reference pose after the field of view adjustment action is completed.

[0091] Optionally, adjusting the field of view of the target image acquisition component back to the specified field of view can restore the parallax geometry of the multi-view vision system, making the optical axes of two adjacent image acquisition components in the multi-view vision system parallel, thereby maximizing the overlapping field of view, which is beneficial for mobile devices to follow and track targets.

[0092] In some embodiments, when the tracked target is lost, the drive adjustment mechanism controls the target image acquisition component (such as a camera that has shifted) to yaw and rotate around the vertical axis or move along a slide rail. By changing the physical orientation of the optical axis of the target image acquisition component, the field of view, which was originally off-center, is made to cover the area where the tracked target may exist. After the target image acquisition component completes the field of view orientation adjustment, motion control commands are generated based on the motion trend of the tracked target before it disappeared or the search results. The mobile device is then controlled to move along the relative direction of the tracked target. During the movement of the mobile device, real-time sampling of each image acquisition component is maintained. Through a target detection network and a re-identification (Re-ID) algorithm, the tracked target is continuously locked in each frame of acquired images, and the tracking data is extracted. The center pixel position of the target's detection box is determined as the target coordinates in the coordinate system of the acquired image. If the target coordinates meet a preset geometric threshold range (for example, if the horizontal / vertical pixel difference between the target coordinates and the center of the acquired image is less than the set threshold, it indicates that the target has been stably located in the center area of ​​the image), the target is identified. If the target is identified by each image acquisition component, the target re-capture is considered successful. At this time, a stop control signal is sent to stop the movement of the mobile device, causing the mobile device to stop moving. At the same time, the control adjustment mechanism restores the field of view of the target image acquisition component that has deflected to the initial posture of looking straight ahead, restoring the standard forward-facing binocular mode.

[0093] For example, when the target is re-detected (e.g., detected in the right camera, i.e., the target image acquisition component is the right camera), a motion control signal is sent to the mobile device, causing the mobile device (such as an electric scooter) to move to the right front until the target reappears within a certain range of another camera (in the current case, the left camera, i.e., the image acquisition component excluding the target image acquisition component) (the center point of the target's detection box is less than a given threshold from the center of the currently acquired image). Then, the forward-facing camera (in the current case, the right camera) is adjusted back to face forward, returning to the standard forward-facing binocular mode.

[0094] In this embodiment, the tracking target is continuously identified from the images acquired by any image acquisition component, and the target coordinates of the tracking target in the images acquired by any image acquisition component are obtained. When the target coordinates meet the preset distance condition, the movement control of the mobile device along the relative direction of the tracking target is stopped, so that the mobile device and the tracking target maintain a certain distance. At the same time, after the movement stops, the field of view of the target image acquisition component is adjusted back to the specified field of view direction, which can eliminate the geometric distortion and parallax accumulation caused by dynamic deflection, quickly restore the standardized observation axis, ensure the orderly connection of field of view retargeting and vehicle displacement in time, and improve the attitude stability and geometric consistency of visual calculation during the target approach phase.

[0095] In an exemplary embodiment, each image acquisition component is respectively disposed on a slide rail; during the process of continuously locating the tracking target based on the sequence of acquired images from the multi-view vision system, the method further includes: controlling at least one image acquisition component in the multi-view vision system to slide along the slide rail based on the distance between the located tracking target and the multi-view vision system, so as to adjust the length of the physical baseline of the multi-view vision system, wherein the length of the physical baseline is positively correlated with the distance between the tracking target and the multi-view vision system.

[0096] In related technologies, cameras typically use a fixed binocular baseline to acquire images. This fixed baseline means that the ranging accuracy based on the left and right visual parallax depends on the distance between the target and the camera. Keeping the baseline constant results in decreased ranging accuracy when the target is too close or too far away. To address this issue, this embodiment dynamically adjusts the sliding motion of at least one image acquisition component in the multi-view vision system along its designated rail based on the distance between the multi-view vision system and the target, thereby improving the ranging accuracy of the multi-view vision system.

[0097] In this embodiment, the length of the physical baseline of the multi-view vision system refers to the actual physical straight-line distance between the optical centers of two adjacent image acquisition components in the multi-view vision system. Optionally, the length of the physical baseline of the multi-view vision system can be a representation of the geometric accuracy and disparity resolution of stereo vision ranging.

[0098] Here, the length of the physical baseline is positively correlated with the distance between the tracked target and the multi-view vision system. It can be understood that the farther the distance between the tracked target and the multi-view vision system, the longer the physical baseline; the closer the distance between the tracked target and the multi-view vision system, the shorter the physical baseline.

[0099] In an optional embodiment, if the distance between the tracked target and the multi-view vision system is greater than a first distance threshold (e.g., 8m), the physical baseline is determined to be a first length (e.g., 100mm); if the distance between the tracked target and the multi-view vision system is less than or equal to the first distance threshold and greater than or equal to a second distance threshold, the physical baseline is determined to be a second length (e.g., 80mm); if the distance between the tracked target and the multi-view vision system is less than the second distance threshold (e.g., 5m), the physical baseline is maintained to a third length (e.g., 60mm), and the movement of the mobile device relative to the tracked target is stopped.

[0100] For example, if the target is re-identified and is at a relatively far distance (e.g., the target is more than 5m away from the mobile device), the position of at least one image acquisition component in the multi-view vision system (e.g., adjusting two image acquisition components (i.e., the left and right cameras)) on the translation rail is controlled according to the distance between the target and the mobile device. If the target is more than 8m away from the mobile device, the length of the physical baseline is adjusted to 100mm; if the target is between 5m and 8m away from the mobile device, the length of the physical baseline is adjusted to 80mm, and the mobile device is controlled to move forward until the distance between the mobile device and the target is less than 5m, at which point the length of the physical baseline is adjusted to 60mm.

[0101] In an optional embodiment, the distance between the tracking target and the multi-view vision system is located by the multi-view vision system, and the movement speed of the tracking target is detected in real time. When the distance between the tracking target and the multi-view vision system is greater than a first distance threshold and the movement speed of the tracking target is greater than a first speed threshold, at least one image acquisition component in the multi-view vision system is adjusted by an adjustment mechanism to slide along its slide rail, so that the physical baseline of the multi-view vision system is a fourth length (e.g., 110 mm). When the distance between the tracking target and the multi-view vision system is less than or equal to the first distance threshold and greater than or equal to a second distance threshold, and the movement speed of the tracking target is less than or equal to the first speed threshold, at least one image acquisition component in the multi-view vision system is adjusted by an adjustment mechanism to slide along its slide rail, so that the physical baseline of the multi-view vision system is a first length. When the distance between the tracking target and the multi-view vision system is less than the second distance threshold and the movement speed of the tracking target is less than or equal to the first speed threshold, at least one image acquisition component in the multi-view vision system is adjusted by an adjustment mechanism to slide along its slide rail, so that the physical baseline of the multi-view vision system is a third length.

[0102] This embodiment establishes a positive correlation between target distance and physical baseline length, and controls the image acquisition component to translate along the slide rail based on the distance between the tracked target and the multi-view vision system. This enables the ranging geometry of the multi-view vision system to adapt to changes in the target's spatial position, effectively overcoming the physical limitations of fixed baseline systems, such as amplified ranging noise at close range due to insufficient parallax and a sharp drop in accuracy at long range due to insufficient parallax.

[0103] In one exemplary embodiment, the multi-view vision system is configured with multiple baseline patterns, each baseline pattern corresponding to a length of a physical baseline, and the multiple baseline patterns corresponding one-to-one with multiple distance ranges; each image acquisition component is respectively disposed on a slide rail; each baseline pattern corresponds to a slide rail position on the slide rail where any image acquisition component is located; based on the distance between the located tracking target and the multi-view vision system, controlling at least one image acquisition component in the multi-view vision system to slide along its respective slide rail includes: determining, among the multiple distance ranges, the target distance range to which the distance between the tracking target and the multi-view vision system belongs, wherein the target distance range corresponds to a target baseline pattern among the multiple baseline patterns; if the current baseline pattern of the multi-view vision system is not the target baseline pattern, controlling each image acquisition component to slide along the slide rail where each image acquisition component is located until it slides to the slide rail position on the slide rail where each image acquisition component is located that corresponds to the target baseline pattern.

[0104] In this embodiment, each of the multiple baseline modes refers to a working mode formed by adjusting the relative positions of the optical centers of two adjacent image acquisition components in a multi-view vision system through a mechanical structure. Here, each of the multiple baseline modes corresponds to a length of the physical baseline. Optionally, the multiple baseline modes may include a standard forward binocular mode (normal tracking mode) (i.e., near baseline mode), a search / re-identification mode (i.e., mid-baseline mode), and a far baseline mode, wherein the physical baseline length corresponding to the standard forward binocular mode (normal tracking mode) is 60 mm, the physical baseline length corresponding to the search / re-identification mode is 80 mm, and the physical baseline length corresponding to the far baseline mode is 100 mm.

[0105] Multiple distance ranges refer to multiple data intervals divided based on the ranging results between the tracked target and the mobile device. For example, multiple distance ranges may include a first distance range (e.g., the distance between the tracked target and the mobile device is greater than 8 meters), a second distance range (e.g., the distance between the tracked target and the mobile device is less than 5 meters), and a third distance range (e.g., the distance between the tracked target and the mobile device is greater than or equal to 5 meters and less than or equal to 8 meters). The first distance range corresponds to the far baseline mode, the second distance range corresponds to the near baseline mode, and the third distance range corresponds to the mid-baseline mode.

[0106] Optionally, a state machine can be set up to drive the switching of the vision mode, that is, to drive the switching of the baseline mode. In the standard forward-facing binocular mode (normal tracking mode), by default, this mode is used for normal communication (the target distance from the scooter is ≤5m). At this time, the camera is in its original position (keeping the baseline as short as possible, the camera facing straight forward, and ensuring overlapping fields of view).

[0107] The target distance range refers to the distance range based on the real-time ranging results between the tracked target and the mobile device. For example, if the real-time ranging result between the tracked target and the mobile device is 6 meters, the target distance range between the tracked target and the multi-view vision system is the third distance range.

[0108] The target baseline pattern refers to the baseline pattern that matches the target distance range. For example, if the target distance between the tracked target and the multi-view vision system belongs to the third distance range, the target baseline pattern is the mid-baseline pattern, and the corresponding physical baseline length of the two adjacent image acquisition components in the multi-view vision system is 80mm.

[0109] In some embodiments, the distance between the tracked target and the mobile device can be obtained through a front-end sensor or binocular parallax calculation. This distance is then compared to a preset distance threshold to determine the corresponding target distance range. Based on a preset mapping relationship between the target distance range and the target baseline mode, the target baseline mode to be switched to is determined. For example, when the distance between the tracked target and the mobile device is determined to be 10m, the corresponding target baseline mode is determined to be the far baseline mode. The currently operating baseline mode is compared with the determined target baseline mode. If the currently operating baseline mode is not the target baseline mode, an adjustment mechanism drives each image acquisition component to slide along its respective slide rail until the position of each image acquisition component on its slide rail matches the position of the slide rail corresponding to the target baseline mode, at which point each image acquisition component stops sliding.

[0110] In this way, when the distance between the target being tracked and the mobile device is far, the physical baseline of the binocular camera can be adjusted by controlling the physical distance between two adjacent image acquisition components (such as the left and right cameras) in the multi-view vision system, thereby improving its ranging accuracy at long distances.

[0111] This embodiment establishes a correspondence between the physical baseline length and the target distance range, enabling the mobile device to automatically adjust the length of the physical baseline when approaching or moving away from the tracking target. This effectively overcomes the problem of distance measurement accuracy attenuation caused by insufficient parallax at close range and inadequate parallax at long range in fixed baseline binocular systems.

[0112] In one exemplary embodiment, the mobile device is an electric scooter, and the multi-view vision system includes two image acquisition units. Two slide rails containing the two image acquisition units are symmetrically arranged on both sides of the handlebar of the electric scooter. Each of the two slide rails contains a slider, and each image acquisition unit is connected to the top of the slider in its respective slide rail via a rotation axis. The electric scooter includes a translation drive motor and a rotation drive motor. The translation drive motor drives the slider in the slide rail containing any image acquisition unit to translate along the slide rail containing any image acquisition unit, and the rotation drive motor drives any image acquisition unit to rotate around the rotation axis of any image acquisition unit.

[0113] In this embodiment, the mobile device can be an electric scooter. The handlebar refers to the horizontal structural component at the front of the electric scooter used to connect the left and right steering control components.

[0114] A slider is a component embedded inside a slide rail, used to drive the image acquisition unit to slide.

[0115] Each image acquisition component is connected to the top of the slider in the slide rail where it is located via a rotation axis. It can be understood that each image acquisition component is mounted on a sliding slider, and the image acquisition component and the slider are connected by a rotation axis, so that the image acquisition component can both move left and right with the slider and rotate independently around its own axis.

[0116] The electric scooter includes a translational drive motor and a rotary drive motor. The translational drive motor is a component that converts electrical energy into linear motion torque, used to drive the slider to slide axially along the rail. Here, the translational drive motor drives the slider within the rail containing each image acquisition component to translate along that rail. The rotary drive motor is a component that converts electrical energy into angular displacement torque, used to drive the image acquisition components to deflect around a rotation axis. Here, the rotary drive motor drives each image acquisition component to rotate around its own rotation axis.

[0117] In this way, when the target is lost or at the edge of the track, the position and direction of the image acquisition component (such as the camera) on the lateral slide rail can be controlled to search for and reposition the target more efficiently, without having to rotate the vehicle body synchronously. This is suitable for the actual constraints of the scooter's low mobility.

[0118] In one embodiment, the multi-view vision system includes a left-eye camera and a right-eye camera. A sliding rail translation mechanism is provided on both sides of the handlebars of the electric scooter. The left-eye and right-eye cameras are respectively installed on both sides of the handlebars (e.g., at the positions of the bells on the left and right handlebars). Horizontal sliding rails and sliders are respectively provided at both ends of the handlebars of the electric scooter. The left-eye and right-eye cameras can translate left and right within a certain range (approximately 2 cm) along the sliding rails. The adjustment mechanism includes a rotation / steering device, where each camera yaws around a vertical axis to adjust the camera's pointing direction in real time, achieving a variable observation orientation. For example, a semi-enclosed aluminum alloy sliding rail is installed on the crossbar at the front of the electric scooter (i.e., on both sides of the handlebars). A high-precision slider is configured inside the sliding rail. The slider integrates a translation drive motor and a camera rotation motor. The camera module is connected to the top of the slider via a rotation axis, allowing it to rotate ±30° around the vertical axis (facing forward → 30° diagonally forward to the left and right respectively). The left and right sliders can translate ±10mm (total travel 20mm) relative to the slide rail, while rotation and sliding are driven by motors. The mounting positions of the left and right cameras are symmetrical relative to the scooter's forehead, with an original baseline (optical center distance between the left and right cameras) of 60mm. Each camera can translate 20mm to either side. Since the intrinsic and extrinsic parameters of the left and right cameras change in real time during dynamic / continuous baseline and attitude focusing, and the maintenance cost of real-time changes is too high for edge-based solutions like scooters, three discrete settings are preset for each camera on the translation slide rail (corresponding to "near baseline," "medium baseline," and "far baseline" modes, with baseline lengths of 60, 80, and 100mm respectively). Precise calibration is performed at each setting, and the mechanical design ensures that the cameras can only be switched to these three settings while limiting mechanical tolerances. In actual use, the corresponding calibration parameters are called for each camera setting to ensure efficient and accurate distance measurement.

[0119] In this embodiment, by symmetrically arranging slide rails and translation drive motors and rotation drive motors on both sides of the handlebar, a dual degree of freedom of lateral translation and rotation about the vertical axis is provided for any image acquisition component, effectively reducing the dependence on the steering ability and maneuverability of the mobile device; the translation and rotation drive motors ensure the accurate execution of baseline reconstruction and field of view deflection actions, improving the response speed and spatial positioning flexibility of the visual tracking hardware platform in complex scenarios.

[0120] An optional embodiment of this application provides a binocular structure adaptive reconstruction method for self-following mobile devices, which involves binocular translation and rotation with three-level variable baseline ranging for electric scooters. Optionally, this embodiment also provides a binocular recognition and ranging device with an independently controllable camera and an adjustable baseline. Based on the position and loss status of the tracked target in the field of view, the baseline and field of view are dynamically adjusted through the translation and rotation of the camera, and combined with state machine control, the visual tracking difficulties caused by scene and physical limitations are solved. Figure 4 This is a schematic diagram of an optional control method according to an embodiment of this application. The specific implementation method is as follows: Figure 4 As shown, it includes:

[0121] Step S41: Locate the tracking target in the acquired image sequence;

[0122] Step S42: Determine if the target being followed has been lost;

[0123] Step S43: Obtain the coordinate position and direction of motion of the target in the historical images to obtain the target's motion trend;

[0124] Step S44: Adjust the field of view of the image acquisition component according to the target movement trend of the tracked target by adjusting the adjustment mechanism;

[0125] Step S45: After rotating the field of view of the target image acquisition component, the tracked target is re-identified.

[0126] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0127] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / random access memory (RAM), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0128] According to another aspect of the embodiments of this application, a mobile device is also provided, which can be used to implement the mobile control method of the mobile device provided in the above embodiments, and will not be repeated hereafter. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0129] Figure 5 This is a structural block diagram of an optional mobile device according to an embodiment of this application, such as... Figure 4 As shown, the mobile device includes a multi-view vision system 502 and a control unit 506. The multi-view vision system includes at least two image acquisition units.

[0130] The multi-view vision system 502 is used for image acquisition.

[0131] Control unit 504 is used to continuously locate a tracking target based on a sequence of acquired images from a multi-view vision system, wherein each acquired image group in the sequence includes at least two acquired images with matching acquisition times; in response to the tracking target being lost, the field of view of the target image acquisition unit is adjusted based on the historical motion information of the tracking target, and the tracking target is searched for from the acquired images of any image acquisition unit in the multi-view vision system, wherein the target image acquisition unit is the image acquisition unit that last identified the tracking target before the tracking target was lost; if the tracking target is found from the acquired images of at least one image acquisition unit in the multi-view vision system, the mobile device is moved along the relative direction of the tracking target.

[0132] It should be noted that the control component 504 in this embodiment can be used to perform the above steps S302 to S306.

[0133] The embodiments provided in this application provide a mobile device that includes a multi-view vision system. The multi-view vision system includes at least two image acquisition components. Based on the historical motion information of the tracked target, the field of view of the target image acquisition component is adjusted to directionally expand the search range. The target is searched from the images acquired by any one of the image acquisition components in the multi-view vision system, achieving efficient local search for the tracked target. When the tracked target is found in the images acquired by at least one image acquisition component in the multi-view vision system, the mobile device is moved along the relative direction of the tracked target. This allows for target re-acquisition with minimal motion cost without increasing the mobile device's own mobility burden, significantly shortening the time required for target re-acquisition and reducing the tracking failure rate caused by the limited mobility of the mobile device. Therefore, the efficiency of the mobile device in searching for and tracking targets can be improved. Thus, this addresses the technical problem in related technologies where the efficiency of mobile devices in searching for and tracking targets is low due to limited field of view and insufficient motion flexibility.

[0134] In an exemplary embodiment, the tracking target is visible within the field of view of the target image acquisition component before it is lost; historical motion information includes motion trend information, which characterizes the target motion trend of the tracking target and is determined based on the historical coordinates of the tracking target identified from the historical images acquired by the target image acquisition component; the control component adjusts the field of view of the target image acquisition component according to the target motion trend so that the change in the field of view of the target image acquisition component is consistent with the target motion trend.

[0135] In an exemplary embodiment, the motion trend information includes a sequence of movement directions. One movement direction in the sequence is represented by the coordinate difference between two historical coordinates of the tracked target identified from two adjacent historical images of the target image acquisition component. Before adjusting the field of view of the target image acquisition component based on the historical motion information of the tracked target, the control component is further configured to acquire the latest N movement directions and the latest historical coordinates in the sequence of movement directions, where N is a positive integer greater than or equal to 2, and the latest historical coordinates are the historical coordinates of the tracked target last identified from the historical images of the target image acquisition component before the current moment. If the distance between the latest historical coordinates and the target image edge in the historical images of the target image acquisition component is less than or equal to a preset distance threshold, a consistency check is performed on the N movement directions. The consistency check is used to check the consistency between the image edge pointed to by the N movement directions and the target image edge. Adjusting the field of view of the target image acquisition component is performed if the consistency check passes.

[0136] In one exemplary embodiment, the control component is further configured to: control the target image acquisition component to rotate based on historical motion information to adjust the field of view direction of the target image acquisition component; control the target image acquisition component to move along a target slide rail based on historical motion information to adjust the field of view range of the target image acquisition component, wherein each image acquisition component is respectively disposed on a slide rail, and the target slide rail is the slide rail where the target image acquisition component is located; wherein, in a multi-view vision system, other image acquisition components besides the target image acquisition component remain stationary relative to the mobile device.

[0137] In an exemplary embodiment, the field of view of the target image acquisition component is adjusted by adjusting the field of view direction of the target image acquisition component; during the process of moving the mobile device along the relative direction of the tracking target, the control component is also used to continuously identify the tracking target from the acquired image of any image acquisition component and obtain the target coordinates of the tracking target in the acquired image of any image acquisition component; when the preset distance condition is met, the movement control of the mobile device along the relative direction of the tracking target is stopped, and the field of view direction of the target image acquisition component is adjusted back to the specified field of view direction.

[0138] In an exemplary embodiment, each image acquisition component is disposed on a slide rail; during the process of continuously locating the tracking target based on the sequence of acquired images from the multi-view vision system, the control component is further configured to control at least one image acquisition component in the multi-view vision system to slide along the slide rail based on the distance between the located tracking target and the multi-view vision system, so as to adjust the length of the physical baseline of the multi-view vision system, wherein the length of the physical baseline is positively correlated with the distance between the tracking target and the multi-view vision system.

[0139] In one exemplary embodiment, the multi-view vision system is configured with multiple baseline patterns, each baseline pattern corresponding to a length of a physical baseline, and the multiple baseline patterns corresponding one-to-one with multiple distance ranges; each image acquisition component is respectively disposed on a slide rail; each baseline pattern corresponds to a slide rail position on the slide rail where the image acquisition component is located; the control component is further configured to determine, within the multiple distance ranges, the target distance range to which the distance between the tracking target and the multi-view vision system belongs, wherein the target distance pattern corresponds to the target baseline pattern among the multiple baseline patterns; if the current baseline pattern of the multi-view vision system is not the target baseline pattern, the control component is configured to slide along the slide rail where the image acquisition component is located until it slides to the slide rail position on the slide rail where the image acquisition component is located that corresponds to the target baseline pattern.

[0140] In one exemplary embodiment, the mobile device is an electric scooter, and the multi-view vision system includes two image acquisition components. Two slide rails containing the two image acquisition components are symmetrically arranged on both sides of the handlebar of the electric scooter. Each of the two slide rails contains a slider, and each image acquisition component is connected to the top of the slider in its respective slide rail via a rotation axis. The electric scooter includes a translation drive motor and a rotation drive motor. The translation drive motor drives the slider in the slide rail containing the image acquisition component to translate along the slide rail containing the image acquisition component, and the rotation drive motor drives the image acquisition component to rotate around the rotation axis of the image acquisition component.

[0141] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.

[0142] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, the computer-readable storage medium including a stored program, wherein the program executes the steps in any of the above method embodiments when it is run.

[0143] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, ROMs, RAMs, portable hard drives, magnetic disks, or optical disks.

[0144] According to another aspect of the embodiments of this application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor is configured to perform the steps of any of the method embodiments described above via the computer program. In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0145] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0146] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0147] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. A method for controlling the movement of a mobile device, characterized in that, Performed by a mobile device, the mobile device including a multi-view vision system, the multi-view vision system including at least two image acquisition components; the method includes: Based on the sequence of images acquired by the multi-view vision system, the tracking target is continuously located, wherein each acquisition image group in the sequence of acquisition images includes at least two acquisition images with matching acquisition times; In response to the tracking target being lost, the field of view of the target image acquisition component is adjusted based on the historical motion information of the tracking target, and the tracking target is searched from the images acquired by any image acquisition component in the multi-view vision system, wherein the target image acquisition component is the image acquisition component that last identified the tracking target before the tracking target was lost; When the tracking target is found in an image acquired from at least one image acquisition component of the multi-view vision system, the mobile device is moved along the relative direction of the tracking target.

2. The method according to claim 1, characterized in that, The historical motion information includes motion trend information, which is used to characterize the target motion trend of the tracked target. The motion trend information is determined based on the historical coordinates of the tracked target identified from the historical images acquired by the target image acquisition component. Adjusting the field of view of the target image acquisition component based on the historical motion information of the tracked target includes: According to the target motion trend, adjust the field of view of the target image acquisition component so that the change in the field of view of the target image acquisition component is consistent with the target motion trend.

3. The method according to claim 2, characterized in that, The motion trend information includes a sequence of movement directions, wherein one of the movement directions in the sequence is represented by the coordinate difference between two historical coordinates of the tracked target identified from two adjacent historical images acquired by the target image acquisition component; Before adjusting the field of view of the target image acquisition component based on the historical motion information of the tracked target, the method further includes: Obtain the latest N motion directions and the latest historical coordinates in the motion direction sequence, where N is a positive integer greater than or equal to 2, and the latest historical coordinates are the historical coordinates of the tracked target last identified from the historical images acquired by the target image acquisition component before the current moment; If the distance between the latest historical coordinates and the target image edge in the historical acquired image of the target image acquisition component is less than or equal to a preset distance threshold, a consistency check is performed on the N motion directions. The consistency check is used to verify the consistency between the image edge pointed to by the N motion directions and the target image edge. Adjusting the field of view of the target image acquisition component is performed if the consistency check passes.

4. The method according to claim 1, characterized in that, Adjusting the field of view of the target image acquisition component based on the historical motion information of the tracked target includes at least one of the following: Based on the historical motion information, the target image acquisition component is controlled to rotate in order to adjust the field of view of the target image acquisition component; Based on the historical motion information, the target image acquisition component is controlled to move along the target slide rail to adjust the field of view of the target image acquisition component. Each image acquisition component is respectively set on a slide rail, and the target slide rail is the slide rail where the target image acquisition component is located. In the multi-view vision system, all image acquisition components other than the target image acquisition component remain stationary relative to the mobile device.

5. The method according to claim 1, characterized in that, The field of view of the target image acquisition component is adjusted by adjusting the field of view direction of the target image acquisition component; During the process of controlling the movement of the mobile device along the relative direction of the tracked target, the method further includes: The tracking target is continuously identified from the images acquired by any of the image acquisition components, and the target coordinates of the tracking target in the images acquired by any of the image acquisition components are obtained; When the preset distance condition is met, the movement control of the mobile device along the relative direction of the tracked target is stopped, and the field of view of the target image acquisition component is adjusted back to the specified field of view.

6. The method according to claim 1, characterized in that, Each of the image acquisition components is respectively mounted on a slide rail; In the process of continuously locating the tracked target based on the sequence of images acquired by the multi-view vision system, the method further includes: Based on the distance between the tracked target and the multi-view vision system, at least one image acquisition component in the multi-view vision system is controlled to slide along its designated slide rail to adjust the length of the physical baseline of the multi-view vision system, wherein the length of the physical baseline is positively correlated with the distance between the tracked target and the multi-view vision system.

7. The method according to claim 6, characterized in that, The multi-view vision system is configured with multiple baseline patterns, each of which corresponds to a length of the physical baseline, and the multiple baseline patterns correspond one-to-one with multiple distance ranges; each image acquisition component is respectively mounted on a slide rail; each baseline pattern corresponds to a slide rail position on the slide rail where the image acquisition component is located; The step of controlling at least one image acquisition component in the multi-view vision system to slide along a slide rail based on the distance between the located tracking target and the multi-view vision system includes: Determine the target distance range to which the distance between the tracked target and the multi-view vision system belongs among the plurality of distance ranges, wherein the target distance range corresponds to the target baseline pattern among the plurality of baseline patterns; If the current baseline mode of the multi-view vision system is not the target baseline mode, control each image acquisition component to slide along the slide rail where the image acquisition component is located until it slides to the slide rail position on the slide rail where the image acquisition component is located that corresponds to the target baseline mode.

8. The method according to claim 6, characterized in that, The mobile device is an electric scooter. The multi-view vision system includes two image acquisition components. The two slide rails containing the two image acquisition components are symmetrically arranged on both sides of the handlebar of the electric scooter. Each of the two slide rails contains a slider. Each image acquisition component is connected to the top of the slider in its respective slide rail via a rotation axis. The electric scooter also includes a translation drive motor and a rotation drive motor. The translation drive motor drives the slider in the slide rail containing the image acquisition component to translate along the slide rail containing the image acquisition component. The rotation drive motor drives the image acquisition component to rotate around its rotation axis.

9. A mobile device, characterized in that, include: A multi-view vision system and control components, wherein the multi-view vision system includes at least two image acquisition components; wherein, The multi-view vision system is used for image acquisition; The control unit is configured to continuously locate the tracked target based on a sequence of acquired images from the multi-view vision system, wherein each acquired image group in the sequence includes at least two acquired images with matching acquisition times; in response to the tracked target being lost, the control unit adjusts the field of view of the target image acquisition unit based on the historical motion information of the tracked target, and searches for the tracked target from the acquired images of any image acquisition unit in the multi-view vision system, wherein the target image acquisition unit is the image acquisition unit that last identified the tracked target before the tracked target was lost; and when the tracked target is found from the acquired images of at least one image acquisition unit in the multi-view vision system, the control unit performs motion control on the mobile device along the relative direction of the tracked target.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.