Control method and electronic equipment
Through communication connection with user equipment, the location, image and sound source data are obtained, and the speaker output is adjusted using neural network models, which solves the problem of inappropriate sound caused by user position changes in tablets and other devices, and improves the user experience.
Patent Information
- Application Number
- CN202510572715.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-18
AI Technical Summary
In electronic devices such as tablets, the speaker sound size is not suitable when users are in different locations, which affects the user experience, especially when the distance is far away or the left and right sides, which leads to inconvenience in use.
By establishing a communication connection with the electronic devices carried by the user, the user's location and image data and sound source data are obtained, and the speaker's output is adjusted using a pre-trained neural network model, accurately locate the user's location and adjust the speaker's sound output.
It realizes real-time adjustment of speaker output according to user location changes, ensuring the appropriate sound size, improving user experience, and reducing positioning errors caused by fluctuations in communication quality.
Smart Images

Figure CN120335757A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and more particularly, to a control method and an electronic device. Background Art
[0002] Electronic products such as tablet computers generally have speakers distributed on the left and right sides. When a user uses a tablet for a voice call, instead of communicating with a handset like a mobile phone, the user receives sound through the tablet's speakers. When the user is far from the tablet or on the left and right sides of the tablet, the sound heard by the user from the speakers is relatively small. If the volume is manually turned up, when the user moves to a position close to the tablet, the sound will be noisy, affecting the user experience. Summary of the Invention
[0003] In view of this, the present disclosure provides a control method and an electronic device.
[0004] A first aspect of the present disclosure provides a control method, the method comprising:
[0005] Obtaining a first position of a target object based on a communication connection with an electronic device carried by the target object;
[0006] Obtaining image data and sound source data of the target object;
[0007] Adjusting the first position based on the image data and the sound source data to obtain a second position of the target object;
[0008] Adjusting the sound output of the speaker based on the second position.
[0009] According to an embodiment of the present disclosure, adjusting the first position based on the image data and the sound source data to obtain a second position of the target object includes:
[0010] Adjusting the direction attribute and the distance attribute of the first position based on at least one of the image data and the sound source data;
[0011] Obtaining the second position according to the adjusted direction attribute and distance attribute.
[0012] According to an embodiment of the present disclosure, adjusting the direction attribute and the distance attribute of the first position based on at least one of the image data and the sound source data includes:
[0013] Inputting the image data and the sound source data into a pre-trained first neural network model respectively to obtain a first confidence level of the image data and a second confidence level of the sound source data; wherein, the first confidence level is related to the number of key points of the target object in the image data; the second confidence level is related to the signal-to-noise ratio of the sound source data;
[0014] Determine a first adjustment value for the direction attribute and a second adjustment value for the distance attribute according to the first confidence level of the image data and the second confidence level of the sound source data;
[0015] Adjust the first position according to the first adjustment value and the second adjustment value to obtain a second position.
[0016] According to an embodiment of the present disclosure, determining a first adjustment value for the direction attribute according to the first confidence level of the image data and the second confidence level of the sound source data includes:
[0017] Determine a first direction and a second direction of the target object according to the image data and the sound source data respectively;
[0018] Determine a first adjustment value for the direction attribute according to the first direction, the second direction, the first confidence level and the second confidence level.
[0019] According to an embodiment of the present disclosure, adjusting the direction attribute and the distance attribute of the first position based on at least one of the image data and the sound source data includes;
[0020] Adjust the distance attribute of the first position according to the image data;
[0021] Adjust the direction attribute of the first position according to the sound source data.
[0022] According to an embodiment of the present disclosure, adjusting the direction attribute and the distance attribute of the first position based on at least one of the image data and the sound source data includes;
[0023] In response to the failure to collect the image data of the target object, adjust the direction attribute and the distance attribute of the first position according to the sound source data.
[0024] According to an embodiment of the present disclosure, adjusting the speaker output based on the second position includes:
[0025] Determine at least one sub-speaker from the speakers according to the direction of the target object represented by the second position;
[0026] In response to the change in the distance of the target object represented by the second position, adjust the output of at least one sub-speaker.
[0027] According to an embodiment of the present disclosure, in response to the change in the distance of the target object represented by the second position, adjusting the output of at least one sub-speaker includes:
[0028] When the distance of the target object represented by the second position increases, increase the output of at least one sub-speaker.
[0029] According to an embodiment of the present disclosure, it further includes:
[0030] After establishing a communication connection, obtain the second position within a historical time period to obtain historical position data;
[0031] Input the historical position data into a pre-trained second neural network model to obtain a prediction result of the target object within a preset time period;
[0032] In response to the interruption of the communication connection, adjust the speaker output based on the prediction result.
[0033] A second aspect of the present disclosure provides an electronic device, including:
[0034] A communication module configured to establish a communication connection with an electronic device carried by a target object and receive the first position of the target object
[0035] A camera and a microphone, respectively configured to obtain image data of the target object and sound source data;
[0036] A processor configured to adjust the first position based on the image data and the sound source data to obtain the second position of the target object; and adjust the sound output of the speaker based on the second position;
[0037] And a speaker.
[0038] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0040] Figure 1 Schematically shows a scenario diagram of a control method according to an embodiment of the present disclosure;
[0041] Figure 2 Schematically shows one of the flowcharts of a control method according to an embodiment of the present disclosure;
[0042] Figure 3 Schematically shows another flowchart of a control method according to an embodiment of the present disclosure;
[0043] Figure 4 Schematically shows a scenario diagram of adjusting the speaker output according to an embodiment of the present disclosure;
[0044] Figure 5 Schematically shows a third flowchart of a control method according to an embodiment of the present disclosure;
[0045] Figure 6 A block diagram of an electronic device according to an embodiment of the present disclosure is schematically shown. Detailed implementation manners
[0046] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0047] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0048] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0049] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).
[0050] Figure 1 A schematic diagram of a scenario of a control method according to an embodiment of the present disclosure is schematically shown.
[0051] As Figure 1 shown, when the position of the user changes within the scene, the user also needs to simultaneously receive the sound from the speaker of the first electronic device 101. If the output of the first electronic device remains unchanged at this time, it may cause the sound effect of the user listening to the external sound of the first electronic device 101 to deteriorate.
[0052] In addition, when the direction or distance of the user relative to the first electronic device 101 changes, the sound effect of receiving the external sound of the first electronic device 101 will also change.
[0053] In an embodiment of the present disclosure, when a user uses a first electronic device 101 to communicate with a remote user, a communication connection with the first electronic device 101 is established through a second electronic device 102 carried by the user himself.
[0054] The method of the embodiment of the present disclosure can be applied to the first electronic device 101. The first electronic device 101 includes, but is not limited to, electronic devices such as tablet computers, laptop computers, and host computers that can be used for external sound. In the present disclosure, only a tablet computer is taken as an exemplary illustration.
[0055] In an embodiment of the present disclosure, the second electronic device 102 includes, but is not limited to, electronic devices such as electronic bracelets, electronic watches, mobile phones, and portable tablets that can be used to establish a communication connection with the first electronic device 101. In the embodiment of the present disclosure, only an electronic watch is taken as an exemplary illustration.
[0056] Figure 2 One of the flowcharts of a control method according to an embodiment of the present disclosure is schematically shown.
[0057] Specifically, as Figure 2 shown, the method includes operations S201 to S204.
[0058] Operation S201: Obtain the first position of the target object based on the communication connection with the electronic device carried by the target object.
[0059] Operation S202: Obtain the image data and sound source data of the target object.
[0060] Operation S203: Adjust the first position based on the image data and sound source data to obtain the second position of the target object.
[0061] Operation S204: Adjust the sound output of the speaker based on the second position.
[0062] Specifically, when the user uses the first electronic device 101, the first electronic device locates the current position of the user based on the communication connection with the second electronic device 102 carried by the user to obtain the first position.
[0063] At the same time, the first electronic device 101 can also obtain the image data and sound source data of the user through hardware devices such as a camera and a microphone. Through the image data and sound source data of the user, the first electronic device 101 can calculate the position information of the position where the user is located, and further use it to adjust the first position, thereby obtaining a more accurate second position of the user.
[0064] When the user moves in the scenario. Based on the above method, the position of the user can be located in real time, and the output of the speaker can be adjusted according to the azimuth and distance from the position to the first electronic device 101, so that even when the position of the user changes, it will not affect the external sound heard by the user from the first electronic device 101.
[0065] By adopting the above method, the first position of the target object relative to the first electronic device 101 is obtained through the communication connection between the two electronic devices, and the image data and sound source data of the target object are obtained through the hardware device of the first electronic device 101 without relying on the communication connection, so that the adjustment process of the first position based on the image data and sound source data is more accurate, avoiding the excessive error of the user's position positioning result caused by poor communication quality.
[0066] Exemplarily, the first electronic device 101 establishes a communication connection with the second electronic device 102 through the ultra-wideband communication method.
[0067] Exemplarily, the first electronic device 101 obtains the UWB data of the user based on the ultra-wideband communication method, and then obtains the first position.
[0068] The ultra-wideband communication method (Ultra Wide Band, UWB) has the advantages of low system complexity, low transmit signal power spectral density, insensitivity to channel fading, low intercept ability, high positioning accuracy, etc., and is especially suitable for high-speed wireless access in indoor and other dense multipath scenarios.
[0069] According to an embodiment of the present disclosure, based on the image data and sound source data, the first position is adjusted to obtain the second position of the target object, including:
[0070] Based on at least one of the image data and sound source data, adjust the direction attribute and distance attribute of the first position; according to the adjusted direction attribute and distance attribute, obtain the second position.
[0071] Specifically, based on the quality of the obtained image data and sound source data, it is determined whether to use one or both of the data to adjust the first position. In the process of adjustment, the second position is obtained by adjusting the direction attribute and distance attribute of the first position.
[0072] Exemplarily, according to the image data and sound source data, the distance attribute of the first position is adjusted; at the same time, according to the image data and sound source data, the direction attribute of the first position is adjusted.
[0073] Figure 3 Schematically shows the second flowchart of a control method according to an embodiment of the present disclosure.
[0074] AsFigure 3 As shown, based on at least one of the image data and the sound source data, the direction attribute and the distance attribute of the first position are adjusted, including operations S301 - S303.
[0075] Operation S301: Input the image data and the sound source data into a pre - trained first neural network model respectively, to obtain a first confidence level of the image data and a second confidence level of the sound source data; wherein, the first confidence level is related to the number of key points of the target object in the image data; the second confidence level is related to the signal - to - noise ratio of the sound source data;
[0076] Operation S302: Determine a first adjustment value of the direction attribute and a second adjustment value of the distance attribute according to the first confidence level of the image data and the second confidence level of the sound source data;
[0077] Operation S303: Adjust the first position according to the first adjustment value and the second adjustment value to obtain a second position.
[0078] In operation S301, a data set is constructed through the image data, the sound source data and the position data obtained based on the communication connection to train the initial neural network model, so as to obtain the pre - trained first neural network model.
[0079] Among them, the first neural network model is based on a deep learning algorithm, can extract the data features of the above - mentioned multiple types of data, and based on the data features, the confidence levels of various types of data can be further determined, including the first confidence level of the image data and the second confidence level of the sound source data. The confidence level represents the credibility of the data.
[0080] The first confidence level of the image data is related to the number of key points of the target object in the image data, and the number of key points is proportional to the first confidence level; the second confidence level is related to the signal - to - noise ratio of the sound source data, and the second confidence level is proportional to the signal - to - noise ratio of the sound source data.
[0081] Exemplarily, the key points obtained by the first electronic device 101 refer to the human body key points captured of the user, such as facial feature points (eyes, nose tip, corners of the mouth, etc.). Based on the key point coordinates, the rotation angles (pitch angle, yaw angle, roll angle) and position coordinates (such as the distance relative to the camera) of the head in the three - dimensional space are calculated. Combining the head pose and the camera calibration parameters of the first electronic device 101, the spatial position (x, y, z) and the orientation angle (such as the angle facing the screen) of the user relative to the first electronic device 101 are output.
[0082] Exemplarily, the signal-to-noise ratio of the sound source data obtained by the first electronic device 101 is the signal-to-noise ratio between the sound emitted by the user and other noises in the environment. When the first electronic device 101 receives the sound source data, it determines the direction and distance of the user through the microphones in different configured directions.
[0083] In the embodiments of the present disclosure, the positional relationship between the user and the first electronic device 101 determined based on the image data is referred to as the image position. The positional relationship between the user and the first electronic device 101 determined based on the sound source data is referred to as the sound position.
[0084] In operation S302, based on the first confidence level and the second confidence level, the first weight and the second weight of the image position and the sound position in the process of adjusting the first position are respectively determined. Among them, the first confidence level is positively correlated with the first weight, and the second confidence level is positively correlated with the second weight.
[0085] In the embodiments of the present disclosure. The third weight characterizes the weight size of the first position in the adjustment process. In the initialization process, the third weight is determined by configuring an initial weight value. As the first electronic device 101 obtains the image data and the sound source data, the initial weight value will change with the communication connection quality between the first electronic device 101 and the second electronic device.
[0086] Exemplarily, the initial weight value is positively correlated with the communication connection quality between the first electronic device 101 and the second electronic device. When the communication connection quality is strong, the weight percentage of the initial weight value W1 increases, and the sum of the weights of the first weight and the second weight W2 decreases. Subsequently, based on the sum of the weights W2, the specific percentages of the first weight and the second weight are respectively determined according to the first confidence level and the second confidence level.
[0087] According to the embodiments of the present disclosure, determining a first adjustment value of the direction attribute according to the first confidence level of the image data and the second confidence level of the sound source data includes:
[0088] According to the image data and the sound source data, the first direction and the second direction of the target object are respectively determined; according to the first direction, the second direction, the first confidence level and the second confidence level, the first adjustment value of the direction attribute is determined.
[0089] Exemplarily, when the sum of the first weight, the second weight, and the third weight is 100%, weights are assigned to the image position and the sound position according to the ratio of the first confidence level and the second confidence level. Further, the product P1 of the first weight and the direction attribute represented by the image position is determined, and the product P2 of the second weight and the direction attribute represented by the sound position is determined. Finally, the first adjustment value P of the direction attribute is obtained as P = P1 + P2. Similarly, the product Q1 of the first weight and the distance attribute represented by the image position is determined, and the product Q2 of the second weight and the distance attribute represented by the sound position is determined. Finally, the second adjustment value Q of the distance attribute is obtained as Q = Q1 + Q2.
[0090] Exemplarily, in operation S303, the direction attribute P and the distance attribute Q of the first position representation are represented as (P, Q); after adjustment, the direction attribute and the distance attribute of the second position representation are represented as (P * W1 + P1 + P2, Q * W1 + Q1 + Q2).
[0091] According to an embodiment of the present disclosure, based on at least one of image data and sound source data, the direction attribute and the distance attribute of the first position are adjusted, including: adjusting the distance attribute of the first position according to the image data; adjusting the direction attribute of the first position according to the sound source data.
[0092] Specifically, when the first electronic device 101 can simultaneously acquire image data and sound source data, the image data includes multiple key points of the user's body captured by the camera. Based on the multiple key points, the change in the direction of the user relative to the first electronic device 101 can be determined more accurately.
[0093] In addition, since the sound source data determines the direction of the user based on the time difference of the sound source data received by multiple microphones in the microphone array of the first electronic device 101. Therefore, when the distance difference between the user's location and each microphone in the microphone array is small, the sound source data cannot well determine the direction of the user at this time, but the distance between the user and the first electronic device 101 can be determined.
[0094] Therefore, based on the data characteristics of the image data and the sound source data, using the image data to adjust the distance attribute of the first position and using the sound source data to adjust the direction attribute of the first position can make the position adjustment structure more accurate.
[0095] According to an embodiment of the present disclosure, based on at least one of image data and sound source data, the direction attribute and the distance attribute of the first position are adjusted, including;
[0096] In response to the failure to collect the image data of the target object, the direction attribute and the distance attribute of the first position are adjusted according to the sound source data.
[0097] Specifically, in some scenarios, since the camera of the first electronic device 101 is blocked, the camera cannot capture the user's image data. At this time, only the sound source data can be used to adjust the direction attribute and distance attribute of the first position, avoiding the inability to complete the position adjustment when the user's image data cannot be obtained.
[0098] According to an embodiment of the present disclosure, adjusting the speaker output based on the second position includes:
[0099] Determining at least one sub-speaker from the speakers according to the direction of the target object characterized by the second position; and adjusting the output of the at least one sub-speaker in response to a change in the distance of the target object characterized by the second position.
[0100] Figure 4 Schematically shows a schematic diagram of a scenario for adjusting the speaker output according to an embodiment of the present disclosure.
[0101] As Figure 4 shown, when the first electronic device includes a plurality of sub-speakers, at least one sub-speaker closest to the direction of the target object characterized by the second position is selected from the plurality of sub-speakers. When it is detected that the distance between the user and the first electronic device 101 changes, the output of at least one sub-speaker is adjusted.
[0102] According to an embodiment of the present disclosure, adjusting the output of at least one sub-speaker in response to a change in the distance of the target object characterized by the second position includes:
[0103] Increasing the output of at least one sub-speaker when the distance of the target object characterized by the second position increases.
[0104] Reducing the output of at least one sub-speaker when the distance of the target object characterized by the second position decreases.
[0105] Exemplarily, while adjusting the output of at least one sub-speaker in response to a change in the distance of the target object characterized by the second position, the output of other sub-speakers is kept unchanged.
[0106] Exemplarily, when the first electronic device includes sub-speakers in two different directions, the angles between the directions of the sub-speakers in the two directions and the direction where the target object is located are the same. At this time, both speakers are configured as at least one sub-speaker.
[0107] Figure 5 Schematically shows a third flowchart of a control method according to an embodiment of the present disclosure.
[0108] According to an embodiment of the present disclosure, operations S501 to S504 are further included.
[0109] Operation S501: After establishing a communication connection, obtain the second position within a historical time period to obtain historical position data;
[0110] Operation S501: Input the historical position data into a pre-trained second neural network model to obtain a prediction result of the target object within a preset time period;
[0111] Operation S501: In response to the interruption of the communication connection, based on the prediction result, adjust the speaker output.
[0112] Specifically, after the first electronic device 101 establishes a communication connection with the second electronic device 102, save the second position within the historical time period, input the historical position data into a pre-trained second neural network model to obtain a prediction result of the target object within a preset time period. According to the prediction result, the movement trajectory of the target object within the preset time period can be predicted, and then the change situation of the position of the target object can be determined.
[0113] When the communication connection is normal, based on the second position, adjust the speaker output. However, due to external interference and other reasons, when the communication quality is poor or the communication connection is interrupted, based on the prediction result, obtain the movement trajectory of the user in the future for a period of time, and then adjust the speaker output.
[0114] By inputting the historical position of the user into the second neural network model to predict the position change of the user in the future for a period of time, it is possible to avoid the situation where the position change of the user cannot be obtained in time due to a sudden interruption of communication. It is ensured that when the first electronic device 101 cannot obtain the first position, it can still adjust the output of the speaker according to the prediction result.
[0115] Next, a detailed process of processing image data and sound source data in the embodiments of the present disclosure will be described.
[0116] Exemplarily, after obtaining the image data, use a convolutional neural network (CNN) to detect key points, such as human body skeleton key points, and then determine the position where the user is located. At the same time, after obtaining the sound source data, extract the sound source direction of the user through beamforming.
[0117] Beamforming is an acoustic processing method that enhances the signal in the target direction and suppresses the noise in the interference direction through spatial filtering technology.
[0118] In the embodiments of the present disclosure, beamforming adjusts the phase and amplitude of the signals of each channel in the microphone array of the first electronic device 101, so that the array forms a "beam" (i.e., signal enhancement) for the sound source in a specific direction, and forms a "null" (i.e., signal suppression) for other directions. At the same time, combined with the voiceprint recognition technology, further distinguish the target voice from the ambient noise to determine the signal-to-noise ratio of the sound source data.
[0119] The process of obtaining the first position through ultra-wideband communication in the embodiments of the present disclosure will be described below.
[0120] Exemplarily, time synchronization is performed between the first electronic device 101 and the second electronic device 102 to ensure the accuracy of timestamps. The transmitting device (such as the first electronic device 101) sends a UWB pulse signal with a timestamp. After the receiving device (the second electronic device 102) receives the signal, it records the reception time and sends a response signal, also with a timestamp. The transmitting device calculates the round-trip time of the signal between the devices based on the timestamps of the transmitted signal and the received response signal. After deducting the signal processing time, the one-way flight time of the signal is obtained. Using the propagation speed (speed of light) of electromagnetic waves and the one-way flight time, the distance relationship between the two devices is calculated. At the same time, the multi-antenna array of the first electronic device 101 measures the angle of arrival (AoA) of the signal, and combines the ranging result to determine the direction relationship between the first electronic device 101 and the second electronic device. By synthesizing the distance relationship and the direction relationship, the first position is obtained.
[0121] The training process of the second neural network model in the embodiments of the present disclosure will be described below.
[0122] Exemplarily, a recurrent neural network (RNN) is used to learn the motion trajectory characterized by the historical change data of the user. The input of the second neural network model is the historical state sequence {X_1, X_2,..., X_k} of the user, and the output is the position prediction result {X_k+1, X_k+2,..., X_k+N} for the next N steps.
[0123] A second aspect of the present disclosure provides an electronic device, including:
[0124] A communication module, configured to establish a communication connection with an electronic device carried by a target object and receive the first position of the target object;
[0125] A camera and a microphone, respectively configured to obtain image data of the target object and sound source data;
[0126] A processor, configured to adjust the first position based on the image data and the sound source data to obtain a second position of the target object; and adjust the sound output of the speaker based on the second position;
[0127] And a speaker.
[0128] The processor is further configured to adjust the direction attribute and the distance attribute of the first position based on at least one of the image data and the sound source data; and obtain the second position according to the adjusted direction attribute and distance attribute.
[0129] The processor is further configured to input the image data and the sound source data into a pre-trained first neural network model respectively, and obtain a first confidence level of the image data and a second confidence level of the sound source data; wherein, the first confidence level is related to the number of key points of the target object in the image data; the second confidence level is related to the signal-to-noise ratio of the sound source data; determine a first adjustment value of the direction attribute and a second adjustment value of the distance attribute according to the first confidence level of the image data and the second confidence level of the sound source data; adjust the first position according to the first adjustment value and the second adjustment value to obtain a second position.
[0130] The processor is further configured to determine a first direction and a second direction of the target object according to the image data and the sound source data respectively; determine a first adjustment value of the direction attribute according to the first direction, the second direction, the first confidence level and the second confidence level.
[0131] The processor is further configured to adjust the distance attribute of the first position according to the image data; adjust the direction attribute of the first position according to the sound source data.
[0132] The processor is further configured to, in response to not collecting the image data of the target object, adjust the direction attribute and the distance attribute of the first position according to the sound source data.
[0133] The processor is further configured to determine at least one sub-speaker from the speakers according to the direction of the target object represented by the second position; in response to a change in the distance of the target object represented by the second position, adjust the output of the at least one sub-speaker.
[0134] The processor is further configured to increase the output of the at least one sub-speaker when the distance of the target object represented by the second position increases.
[0135] The processor is further configured to: after establishing the communication connection, obtain the second position in the historical time period to obtain historical position data; input the historical position data into a pre-trained second neural network model to obtain a prediction result of the target object within a preset time period; in response to the interruption of the communication connection, adjust the speaker output based on the prediction result.
[0136] It should be noted that the electronic device part in the embodiments of the present disclosure corresponds to the processing method part in the embodiments of the present disclosure, and their specific implementation details are also the same, which will not be elaborated here.
[0137] Figure 6 A block diagram of another electronic device suitable for implementing the method described above according to an embodiment of the present disclosure is schematically shown. Figure 6The illustrated electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0138] As Figure 6 shown, the electronic device 600 according to an embodiment of the present disclosure includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as, for example, an application specific integrated circuit (ASIC)), and so on. The processor 601 can also include on-board memory configured for caching purposes. The processor 601 can include a single processing unit or multiple processing units configured to perform different actions of the method flow according to an embodiment of the present disclosure.
[0139] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to an embodiment of the present disclosure by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs can also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 can also perform various operations of the method flow according to an embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0140] According to an embodiment of the present disclosure, the electronic device 600 can further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 can further include one or more of the following components connected to the input / output (I / O) interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read therefrom can be installed into the storage section 608 as needed.
[0141] According to an embodiment of the present disclosure, the method flow according to the embodiments of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program code configured to execute the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 609, and / or installed from a removable medium 611. When the computer program is executed by a processor 601, the above functions defined in the system according to the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described system, device, apparatus, module, unit, etc. can be implemented by computer program modules.
[0142] The present disclosure also provides a computer-readable storage medium, which can be included in the device / device / system described in the above embodiments; or can exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.
[0143] According to an embodiment of the present disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. For example, it can include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, device, or device.
[0144] For example, according to an embodiment of the present disclosure, the computer-readable storage medium can include the above-described ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603.
[0145] An embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program includes program code configured to execute the method provided by the embodiments of the present disclosure. When the computer program product runs on an electronic device, the program code is configured to enable the electronic device to implement the remote sensing image detection method based on a deep neural network provided by the embodiments of the present disclosure.
[0146] When the computer program is executed by the processor 601, the above functions defined in the system / apparatus of the embodiments of the present disclosure are executed. According to the embodiments of the present disclosure, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0147] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program can also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 609, and / or be installed from the removable medium 611. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0148] According to the embodiments of the present disclosure, the program code configured to execute the computer program provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include but are not limited to such as Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0149] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions configured to implement the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in various embodiments of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0150] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A control method, the method comprising: Obtaining a first position of the target object based on a communication connection with an electronic device carried by the target object; Obtaining image data and sound source data of the target object; Adjusting the first position based on the image data and the sound source data to obtain a second position of the target object; Adjusting the sound output of a speaker based on the second position.
2. The method according to claim 1, wherein adjusting the first position based on the image data and the sound source data to obtain a second position of the target object comprises: Adjusting a direction attribute and a distance attribute of the first position based on at least one of the image data and the sound source data; Obtaining the second position according to the adjusted direction attribute and distance attribute.
3. The method according to claim 2, wherein adjusting a direction attribute and a distance attribute of the first position based on at least one of the image data and the sound source data comprises: Inputting the image data and the sound source data into a pre-trained first neural network model respectively to obtain a first confidence level of the image data and a second confidence level of the sound source data; wherein, the first confidence level is related to the number of key points of the target object in the image data; the second confidence level is related to the signal-to-noise ratio of the sound source data; Determining a first adjustment value of the direction attribute and a second adjustment value of the distance attribute according to the first confidence level of the image data and the second confidence level of the sound source data; Adjusting the first position according to the first adjustment value and the second adjustment value to obtain the second position.
4. The method according to claim 3, wherein determining a first adjustment value of the direction attribute according to the first confidence level of the image data and the second confidence level of the sound source data comprises: Determining a first direction and a second direction of the target object according to the image data and the sound source data respectively; Determining a first adjustment value of the direction attribute according to the first direction, the second direction, the first confidence level and the second confidence level.
5. The method according to claim 2, wherein adjusting a direction attribute and a distance attribute of the first position based on at least one of the image data and the sound source data comprises; Adjusting the distance attribute of the first position according to the image data; Adjusting the direction attribute of the first position according to the sound source data.
6. The method according to claim 2, wherein adjusting a direction attribute and a distance attribute of the first position based on at least one of the image data and the sound source data comprises; In response to not collecting the image data of the target object, adjusting the direction attribute and the distance attribute of the first position according to the sound source data.
7. The method according to claim 1, wherein adjusting the speaker output based on the second position comprises: Determining at least one sub-speaker from the speakers according to the direction of the target object represented by the second position; Adjust the output of the at least one sub-speaker in response to a change in the distance of the target object represented by the second location.
8. The method according to claim 7, wherein adjusting the output of the at least one sub-speaker in response to a change in the distance of the target object represented by the second location comprises: Increasing the output of the at least one sub-speaker when the distance of the target object represented by the second location increases.
9. The method according to claim 1, further comprising: After establishing the communication connection, obtaining the second location within a historical time period to obtain historical location data; Inputting the historical location data into a pre-trained second neural network model to obtain a prediction result of the target object within a preset time period; In response to the communication connection being interrupted, adjusting the speaker output based on the prediction result.
10. An electronic device, comprising: A communication module configured to establish a communication connection with an electronic device carried by a target object and receive a first location of the target object; A camera and a microphone respectively configured to obtain image data of the target object and sound source data; A processor configured to adjust the first location based on the image data and the sound source data to obtain a second location of the target object; and adjust the sound output of a speaker based on the second location; And, the speaker.