Decoding method and electronic device
By referring to the state parameters of the previous moment in the electronic device at a subsequent moment in the electronic device and updating the state parameters to discard irrelevant semantic features, the problem of text inconsistency caused by semantic vector independence in bidirectional decoding is solved, and the semantic coherence and accuracy of text output is improved.
Patent Information
- Application Number
- CN202011182533.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-10-29
AI Technical Summary
In the prior art, multiple sets of semantic vectors obtained by electronic devices during bidirectional decoding are relatively independent, resulting in semantic incoherence of the output text and reducing accuracy.
By acquiring the data group at the first moment and decoding it in both directions, the status parameters are output, and then referring to the status parameters at the previous moment when acquiring the data group at the subsequent moment, the status parameters are updated to discard irrelevant semantic features to ensure semantic coherence.
The semantic coherence and accuracy of text output are improved, and the consistency between semantic vector groups is ensured by referring to the semantic features of the previous moment at a subsequent moment, thereby improving the accuracy of text.
Smart Images

Figure CN114512138B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of signal decoding, and in particular, to a decoding method and an electronic device. Background Art
[0002] In application scenarios such as using a voice assistant or image scanning, an electronic device needs to convert an analog signal (such as a voice signal or an image signal containing text information) input by a user into text. For example, the electronic device can collect voice frames input by the user and extract the sound features (such as pitch, loudness, or timbre, etc.) of each voice frame in real time to obtain a frame vector corresponding to the voice frame. Furthermore, the electronic device can decode the frame vector of each voice frame to obtain a corresponding semantic vector, and this semantic vector is used to indicate the semantic features (such as part of speech, emotional color, etc.) of the corresponding voice frame. Then, the electronic device inputs the semantic vector into a text generation model to output the text corresponding to the semantic vector.
[0003] In some scenarios, after processing a voice signal or an image signal into multiple frame vectors, the electronic device can use a bidirectional decoder to decode the frame vectors. The number of frame vectors that can be decoded in one decoding process is preset in the bidirectional decoder. Taking the number as 2 as an example, as Figure 1 shown, after the bidirectional decoder obtains the frame vector X1 and the frame vector X2, the forward sequence of these two frame vectors is: X1, X2, and the reverse sequence of these two frame vectors is: X2, X1. Furthermore, the electronic device decodes the above forward sequence and reverse sequence simultaneously. For example, it can decode X1 in the forward sequence based on a preset initial forward state parameter FS0, and output a first forward state parameter FS1 and a forward decoding result y1. At the same time, it decodes X2 in the reverse sequence based on a preset initial reverse state parameter BS0, and outputs a first reverse state parameter BS1 and a reverse decoding result y3. Then, based on y1 and y3, the electronic device can determine the semantic vector Y1 corresponding to the frame vector X1. That is to say, when determining the semantic vector Y1, the reverse decoding result y3 of X2 and the preset initial forward state parameter FS0 are referred to, so that the accuracy of the semantic vector Y1 is higher. Similarly, X2 in the forward sequence can be decoded according to the first forward state parameter FS1 to output a forward decoding result y2, and at the same time, X1 in the reverse sequence can be decoded according to the first reverse state parameter BS1 to output a reverse decoding result y4. Then, based on y2 and y4, the electronic device can determine the semantic vector Y2 corresponding to the frame vector X2. Similarly, when determining the semantic vector Y2, the first forward state parameter FS1 and the reverse decoding result y4 of X1 are referred to, so that the accuracy of the semantic vector Y2 is higher.
[0004] Subsequently, when the bidirectional decoder obtains the frame vectors X3 and X4, the electronic device can re-decode according to the above method to obtain the corresponding semantic vectors Y3 and Y4. It can be seen that the first set of semantic vectors (such as semantic vectors Y1 and Y2) obtained by the electronic device is relatively independent of the second set of semantic vectors (such as semantic vectors Y3 and Y4), which may cause the semantics of the text determined by the electronic device using this decoding method to be incoherent, resulting in a decrease in the accuracy of the finally output text. Summary of the Invention
[0005] An embodiment of the present application provides a decoding method and an electronic device, which can solve the problem that multiple sets of obtained semantic vectors are relatively independent, thereby improving the coherence and accuracy of the semantics of the determined text.
[0006] To achieve the above object, the present application adopts the following technical solutions:
[0007] In a first aspect, a decoding method is provided. The decoding method includes: obtaining a first data set at a first moment. The first data set includes N data frames, where N is an integer greater than 1. Bidirectionally decoding the first data set to output a first state parameter and semantic vectors corresponding to each of the N data frames. The first state parameter is used to indicate the semantic feature of the first data set. Obtaining a second data set at a second moment. The second data set includes M data frames, the second moment is later than the first moment, and M is an integer greater than 1. Bidirectionally decoding the second data set according to the first state parameter to output semantic vectors corresponding to each of the M data frames. Thus, when bidirectionally decoding the second data set, the first state parameter is referred to, that is, the semantic feature of the first data set is referred to, so that the semantic vector group corresponding to the first data set and the semantic vector group corresponding to the second data set are coherent, thereby improving the coherence and accuracy of the semantics of the determined text.
[0008] In a possible design approach, the first state parameter may include a first forward state parameter and a first reverse state parameter. Among them, bidirectionally decoding the first data group and outputting the first state parameter may include: unidirectionally decoding the forward sequence of the first data group to output the first forward state parameter. The first forward state parameter is used to indicate the semantic features of the forward sequence of the first data group and the semantic features of the data frame located at the end of the forward sequence of the first data group. The forward sequence of the first data group is used to indicate the frame sequence obtained by sorting N data frames according to the timing relationship of the acquired N data frames. Unidirectionally decoding the reverse sequence of the first data group to output the first reverse state parameter. The first reverse state parameter is used to indicate the semantic features of the reverse sequence of the first data group and the semantic features of the data frame located at the end of the reverse sequence of the first data group. The reverse sequence of the first data group is used to indicate the frame sequence obtained by sorting N data frames according to the reverse timing relationship of the acquired N data frames.
[0009] Further, bidirectionally decoding the second data group according to the first state parameter may include: updating the first state parameter to a second state parameter, and the second state parameter discards the semantic features of the reverse sequence of the first data group; bidirectionally decoding the second data group according to the second state parameter. It can be understood that the semantic features of the forward sequence of the first data group can express the semantics actually input by the user at the first moment, and have high reliability as a reference basis for bidirectionally decoding the second data group, so they are retained; while the semantic features of the reverse sequence of the first data group cannot express the semantics actually input by the user at the first moment, and have low reliability as a reference basis for bidirectionally decoding the second data group, so they are discarded. Thus, the second state parameter has high reliability as a reference basis for bidirectionally decoding the second data group.
[0010] Furthermore, the second state parameter may include a second forward state parameter and a second reverse state parameter, and the second forward state parameter is equal to the second reverse state parameter. The activation layer parameter in the second forward state parameter is used to indicate: the semantic features of the forward sequence of the first data group. The hidden layer parameter in the second forward state parameter is used to indicate: the semantic features of the data frame located at the end of the forward sequence of the first data group and the semantic features of the data frame located at the end of the reverse sequence of the first data group. The semantic features of the forward sequence of the first data group can express the semantics actually input by the user at the first moment, so it has high reliability as a reference basis for bidirectionally decoding the second data group; in addition, the semantic features of the data frame located at the end of the forward sequence of the first data group and the semantic features of the data frame located at the end of the reverse sequence of the first data group are the reference basis for generating the semantic vector corresponding to the data frame located at the end of the first data group, so it has high reliability as a reference basis for bidirectionally decoding the second data group.
[0011] Furthermore, the first hidden layer parameter in the first forward state parameter can be a first matrix; the second hidden layer parameter in the first reverse state parameter can be a second matrix. The hidden layer parameter in the second forward state parameter is a matrix formed by concatenating or stacking the first matrix and the second matrix.
[0012] In a possible design, the first forward state parameter may include a first hidden layer parameter and a first activation layer parameter. The first hidden layer parameter is used to indicate the semantic features of the data frame at the end of the forward sequence of the first data group, and the first activation layer parameter is used to indicate the semantic features of the forward sequence of the first data group. The first reverse state parameter includes a second hidden layer parameter and a second activation layer parameter. The second hidden layer parameter is used to indicate the semantic features of the data frame at the end of the forward sequence of the first data group, and the second activation layer parameter is used to indicate the semantic features of the forward sequence of the first data group.
[0013] In a possible design, the method may further include: while outputting the semantic vectors corresponding to each of the M data frames, outputting a third state parameter. The third state parameter is used to indicate the semantic features of the second data group. Obtain a third data group at a third moment. The third data group includes L data frames. The third moment is later than the second moment, and L is an integer greater than 1. Bidirectionally decode the third data group according to the third state parameter, and output the semantic vectors corresponding to each of the L data frames. Thus, when bidirectionally decoding the third data group, the third state parameter is referred to, that is, the semantic features of the second data group are referred to, so that the semantic vector groups corresponding to the second data group and the third data group are coherent, thereby increasing the length of the text with determined semantic coherence and accuracy.
[0014] Furthermore, the third state parameter may include a third forward state parameter and a third reverse state parameter. Among them, bidirectionally decoding the third data group and outputting the third state parameter includes: unidirectionally decoding the forward sequence of the third data group and outputting the third forward state parameter. The third forward state parameter is used to indicate the semantic features of the forward sequence of the third data group and the semantic features of the data frame at the end of the forward sequence of the third data group. The forward sequence of the third data group is used to indicate the frame sequence obtained by sorting the L data frames according to the temporal relationship of the obtained L data frames. Unidirectionally decode the reverse sequence of the third data group and output the third reverse state parameter. The third reverse state parameter is used to indicate the semantic features of the reverse sequence of the third data group and the semantic features of the data frame at the end of the reverse sequence of the third data group. The reverse sequence of the third data group is used to indicate the frame sequence obtained by sorting the L data frames according to the reverse temporal relationship of the obtained L data frames.
[0015] Further, bidirectionally decoding the third data group according to the third state parameter may include: updating the third state parameter to a fourth state parameter, where the fourth state parameter discards the semantic features of the reverse sequence of the second data group. Bidirectionally decoding the third data group according to the fourth state parameter. It can be understood that the semantic features of the forward sequence of the second data group can express the semantics actually input by the user at the second moment, and are highly reliable as a reference basis for bidirectionally decoding the third data group, so they are retained; while the semantic features of the reverse sequence of the second data group cannot express the semantics actually input by the user at the second moment, and are not highly reliable as a reference basis for bidirectionally decoding the third data group, so they are discarded. Thus, the fourth state parameter is highly reliable as a reference basis for bidirectionally decoding the third data group.
[0016] In a possible design, bidirectionally decoding the first data group may include: in the first decoder, bidirectionally decoding the first data group. Bidirectionally decoding the second data group according to the first state parameter may include: in the second decoder, bidirectionally decoding the second data group according to the first state parameter. It can be understood that when bidirectionally decoding the second data group at the second moment, the first data group has been decoded. Thus, at the second moment, the first decoder can obtain the third data group and bidirectionally decode it, thereby improving the decoding efficiency.
[0017] In a possible design, all N data frames are speech frames. Obtaining the first data group at the first moment may include: in response to an input speech signal, extracting speech frames from the speech signal. When N speech frames are extracted at the first moment, the N speech frames form the first data group. Alternatively, all N data frames are image frames. Obtaining the first data group at the first moment may include: in response to an input image signal, extracting image frames from the image signal. When N image frames are extracted at the first moment, the N image frames form the first data group.
[0018] In a possible design, the semantic features of the first data group are associated with the semantic features of the second data group.
[0019] In a second aspect, an embodiment of the present application further provides an electronic device, including:
[0020] One or more processors;
[0021] A memory;
[0022] One or more processors;
[0023] A memory;
[0024] And one or more computer programs, where one or more computer programs are stored in the memory, and when the computer programs are executed by one or more processors, the electronic device is caused to execute the decoding method as described in any one of the above first aspects.
[0025] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium, which includes a computer program or instruction. When the computer program or instruction runs on a computer, the computer is caused to execute the decoding method described in any one of the above first aspects.
[0026] In a fourth aspect, an embodiment of the present application further provides a computer program product, which includes: a computer program or instruction. When the computer program or instruction runs on a computer, the computer is caused to execute the decoding method described in any one of the above first aspects.
[0027] It can be understood that the electronic device provided in the above second aspect, the computer-readable storage medium provided in the third aspect, and the computer program product provided in the fourth aspect are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 Principle of bidirectional decoding provided by the prior art Figure 1 ;
[0029] Figure 2 Schematic diagram of the principle of unidirectional decoding provided by the prior art;
[0030] Figure 3 Schematic diagram of the hardware structure of the mobile phone provided by an embodiment of the present application;
[0031] Figure 4 Flow of the decoding method provided by an embodiment of the present application Figure 1 ;
[0032] Figure 5 Principle of bidirectional decoding provided by an embodiment of the present application Figure 2 ;
[0033] Figure 6 Flow of the decoding method provided by an embodiment of the present application Figure 2 ;
[0034] Figure 7 Schematic diagram of the software structure of the decoding device provided by an embodiment of the present application;
[0035] Figure 8 Schematic diagram of the structural composition of the electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The following introduces the technical terms involved in the embodiments of the present application.
[0037] Unidirectional decoding: AsFigure 2 As shown, when the electronic device obtains the frame vector X1 at the first moment, it can decode the frame vector X1 according to the pre-configured initial state parameter FS0, and output the decoding result Y1 and the first state parameter FS1. Among them, FS1 represents the semantic feature corresponding to the frame vector X1. When the electronic device obtains the frame vector X2 at the second moment, it can decode the frame vector X2 according to the first state parameter FS1, and output the decoding result Y2 and the second state parameter FS2. Among them, FS2 represents the semantic feature corresponding to the frame vector X2. Similarly, when the electronic device obtains the frame vector X3 at the third moment, it can decode the frame vector X3 according to the second state parameter FS2 according to the above method, and output the decoding result Y3.
[0038] It can be seen that in the above one-way decoding process, in the electronic device, the semantic feature corresponding to the frame vector X1 is input to decode the frame vector X2; in the electronic device, the semantic feature corresponding to the frame vector X2 is input to decode the frame vector X3. Therefore, in the one-way decoding process, for the decoding of each frame vector, only the semantic feature of the previous frame vector of the frame vector to be decoded is referred to.
[0039] Bidirectional decoding: The number of frame vectors that can be decoded in one decoding process is preset in the electronic device. Taking the number as 2 as an example, as Figure 2 shown, after the electronic device obtains the frame vector X1 and the frame vector X2, the forward sequence of these two frame vectors is: X1, X2, and the reverse sequence of these two frame vectors is: X2, X1. Furthermore, the electronic device decodes the above forward sequence and reverse sequence simultaneously. For example, the electronic device can decode X1 in the forward sequence based on the pre-set initial forward state parameter FS0, and output the first forward state parameter FS1 and the forward decoding result y1. Among them, FS1 represents the semantic feature corresponding to the frame vector X1 in the forward sequence. At the same time, the electronic device decodes X2 in the reverse sequence based on the pre-set reverse state parameter BS0, and outputs the first reverse state parameter BS1 and the reverse decoding result y3. BS1 represents the semantic feature corresponding to the frame vector X2 in the reverse sequence.
[0040] Then, based on y1 and y3, the electronic device can determine the semantic vector Y1 corresponding to the frame vector X1. That is to say, when determining the semantic vector Y1, the reverse decoding result y3 of X2 and the pre-set initial forward state parameter FS0 are referred to, so that the accuracy of the semantic vector Y1 is higher. Similarly, the electronic device can decode X2 in the forward sequence according to the first forward state parameter FS1, and output the forward decoding result y2. At the same time, the electronic device decodes X1 in the reverse sequence according to the first reverse state parameter BS1, and outputs the reverse decoding result y4.
[0041] Next, the technical solutions in the present application will be described with reference to the accompanying drawings.
[0042] A decoding method provided by an embodiment of the present application can be applied to an electronic device, which can be a mobile phone, a tablet computer, a notebook computer, an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a personal digital assistant (PDA), a wearable electronic device, a virtual reality device, etc. The embodiments of the present application do not impose any restrictions on this.
[0043] Exemplarily, as Figure 3 shown, the electronic device in the embodiment of the present application can be a mobile phone 300. The following takes the mobile phone 300 as an example to specifically illustrate the embodiment. It should be understood that the illustrated mobile phone 300 is only an example of the above-mentioned electronic device, and the mobile phone 300 may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations.
[0044] As Figure 3 shown, the mobile phone 300 may specifically include: a processor 301, and a radio frequency (RF) circuit 302, a memory 303, a touch screen 304, a Bluetooth device 305, one or more sensors 306, a wireless fidelity (Wi-Fi) device 307, a positioning device 308, an audio circuit 309, a peripheral interface 320, a camera 321, and a power system 111, etc., which are respectively electrically connected to the processor 301. These components can communicate through one or more communication buses or signal lines ( Figure 3 not shown in the figure). Those skilled in the art can understand that Figure 3 the hardware structure shown in the figure does not constitute a limitation on the mobile phone. The mobile phone 300 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0045] The following Figure 3 specifically introduces each component of the mobile phone 300:
[0046] The processor 301 is the control center of the mobile phone 300, connecting various parts of the mobile phone 300 through various interfaces and lines. By running or executing applications stored in the memory 303 and calling data stored in the memory 303, it executes various functions of the mobile phone 300 and processes data. In some embodiments of the present application, the above-mentioned processor 301 may further include a fingerprint verification chip for verifying the collected fingerprints.
[0047] The radio frequency circuit 302 can be used for receiving and sending wireless signals during information transceiver or call processes. In particular, the radio frequency circuit 302 can receive the downlink data from the base station and process it by the processor 301; in addition, it can send the data related to the uplink to the base station. Generally, the radio frequency circuit includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, etc. In addition, the radio frequency circuit 302 can also communicate with other devices through wireless communication. The wireless communication can use any wireless communication standard or protocol, including but not limited to Global System for Mobile Communications, General Packet Radio Service, Code Division Multiple Access, Wideband Code Division Multiple Access, Long Term Evolution, email, Short Message Service, etc.
[0048] The memory 303 is used to store applications and data. The processor 301 executes various functions and data processing of the mobile phone 300 by running the applications and data stored in the memory 303. The memory 303 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and applications required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store the data created when using the mobile phone 300 (such as audio data, a phone book, etc.). In addition, the memory 303 can include a high-speed random access memory (RAM), and can also include a non-volatile memory, such as a disk storage device, a flash memory device, or other volatile solid-state storage devices, etc. The memory 303 can store various operating systems. For example, the operating system developed by Apple Inc., the operating system developed by Google Inc., etc. The above-mentioned memory 303 can be independent and connected to the processor 301 through the above-mentioned communication bus; the memory 303 can also be integrated with the processor 301.
[0049] The touch screen 304 can specifically include a touchpad 304-1 and a display 304-2.
[0050] Among them, the touchpad 304-1 can collect the touch events of the user of the mobile phone 300 on or near it (such as the operations of the user using a finger, a stylus, or any suitable object on or near the touchpad 304-1), and send the collected touch information to other devices (such as the processor 301). Among them, the touch events of the user near the touchpad 304-1 can be called hovering touch; hovering touch can mean that the user does not need to directly touch the touchpad to select, move, or drag a target (such as a control, etc.), but only needs the user to be near the terminal to perform the desired function. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touchpad 304-1.
[0051] The display (also known as the display screen) 304-2 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone 300. The display 304-2 can be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The touchpad 304-1 can be covered on the display 304-2. After the touchpad 304-1 detects a touch event on or near it, it is transmitted to the processor 301 to determine the type of the touch event. Subsequently, the processor 301 can provide a corresponding visual output on the display 304-2 according to the type of the touch event. Although in Figure 3 in, the touchpad 304-1 and the display screen 304-2 are implemented as two independent components to realize the input and output functions of the mobile phone 300, but in some embodiments, the touchpad 304-1 and the display screen 304-2 can be integrated to realize the input and output functions of the mobile phone 300. It can be understood that the touch screen 304 is stacked by multiple layers of materials. In the embodiments of the present application, only the touchpad (layer) and the display screen (layer) are shown, and other layers are not described in the embodiments of the present application. In addition, the touchpad 304-1 can be configured on the front of the mobile phone 300 in the form of a full panel, and the display screen 304-2 can also be configured on the front of the mobile phone 300 in the form of a full panel, so that a borderless structure can be realized on the front of the mobile phone 300.
[0052] In addition, the mobile phone 300 can also have a fingerprint recognition function. For example, a fingerprint recognizer 312 can be configured on the back of the mobile phone 300 (such as below the rear camera), or a fingerprint recognizer 312 can be configured on the front of the mobile phone 300 (such as below the touch screen 304). For another example, a fingerprint acquisition device 312 can be configured in the touch screen 304 to realize the fingerprint recognition function, that is, the fingerprint acquisition device 312 can be integrated with the touch screen 304 to realize the fingerprint recognition function of the mobile phone 300. In this case, the fingerprint acquisition device 312 is configured in the touch screen 304, and can be a part of the touch screen 304 or configured in the touch screen 304 in other ways. The main component of the fingerprint acquisition device 312 in the embodiments of the present application is a fingerprint sensor, and the fingerprint sensor can adopt any type of sensing technology, including but not limited to optical, capacitive, piezoelectric or ultrasonic sensing technology, etc.
[0053] The mobile phone 300 can also include a Bluetooth device 305, which is used to realize data exchange (such as, transceiver translation text and original text) between the mobile phone 300 and other short-range terminals (such as mobile phones, smart watches, etc.). The Bluetooth device 305 in the embodiments of the present application can be an integrated circuit or a Bluetooth chip, etc.
[0054] The mobile phone 300 may also include at least one sensor 306, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display of the touch screen 304 according to the brightness of the ambient light, and the proximity sensor can turn off the power of the display when the mobile phone 300 is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used in applications for identifying the posture of the mobile phone (such as landscape / portrait screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the mobile phone 300 can also be configured with, they will not be elaborated here.
[0055] A Wi-Fi device 307 is used to provide network access for the mobile phone 300 that complies with Wi-Fi related standard protocols. The mobile phone 300 can access a Wi-Fi access point through the Wi-Fi device 307, thereby helping the user to send and receive information, and at the same time providing the user with wireless broadband Internet access. In some other embodiments, the Wi-Fi device 307 can also act as a Wi-Fi wireless access point and can provide Wi-Fi network access for other terminals.
[0056] A positioning device 308 is used to provide the geographical location for the mobile phone 300. It can be understood that the positioning device 308 can specifically be a receiver of a global positioning system (GPS), or a Beidou satellite navigation system, a Russian GLONASS and other positioning systems. After receiving the geographical location sent by the above positioning system, the positioning device 308 sends this information to the processor 301 for processing, or sends it to the memory 303 for storage. In some other embodiments, the positioning device 308 can also be a receiver of an assisted global positioning system (AGPS). The AGPS system assists the positioning device 308 to complete ranging and positioning services by acting as an auxiliary server. In this case, the auxiliary positioning server communicates with the positioning device 308 (i.e., the GPS receiver) of the terminal such as the mobile phone 300 through a wireless communication network to provide positioning assistance. In some other embodiments, the positioning device 308 can also be a positioning technology based on Wi-Fi access points. Since each Wi-Fi access point has a globally unique media access control (MAC) address, the terminal can scan and collect the broadcast signals of the surrounding Wi-Fi access points when the Wi-Fi is turned on, so that the MAC address broadcast by the Wi-Fi access point can be obtained; the terminal sends the data (such as the MAC address) that can identify the Wi-Fi access point to the location server through a wireless communication network. The location server retrieves the geographical location of each Wi-Fi access point, and calculates the geographical location of the terminal in combination with the strength of the Wi-Fi broadcast signal and sends it to the positioning device 308 of the terminal.
[0057] The audio circuit 309, the speaker 313, and the microphone 314 can provide an audio interface between the user and the mobile phone 300. The audio circuit 309 can transmit the electrical signal converted from the received audio data to the speaker 313, and the speaker 313 converts it into a sound signal for output; on the other hand, the microphone 314 can convert the collected sound signal into an electrical signal and output the electrical signal to the audio circuit 309. The audio circuit 309 converts the electrical signal into audio data. Furthermore, the audio circuit 309 can output the audio data to the memory 303 so that the processor 301 can further process the audio data into speech frames.
[0058] The peripheral interface 320 is used to provide various interfaces for external input / output devices (such as keyboards, mice, external displays, external memories, subscriber identification module cards, etc.). For example, it is connected to a mouse through a universal serial bus (USB) interface and connected to a subscriber identification module (SIM) card provided by a telecommunications operator through the metal contacts on the subscriber identification module card slot. The peripheral interface 320 can be used to couple the above-mentioned external input / output peripheral devices to the processor 301 and the memory 303.
[0059] The camera 321 can input the collected image signal into the processor 301, and the processor 301 can process the image signal into an image frame.
[0060] The mobile phone 300 may further include a power supply device 311 (such as a battery and a power management chip) for powering each component. The battery can be logically connected to the processor 301 through the power management chip, so as to realize functions such as management of charging, discharging, and power consumption management through the power supply device 311.
[0061] It should be noted that the solution in the embodiment of the present application can also be applied to other electronic devices, and the corresponding names can also be replaced with the names of the corresponding functions in other electronic devices.
[0062] Exemplarily, Figure 4 FIG. is a schematic flowchart of the decoding method provided by the embodiment of the present application. The decoding method can be applied to the above-mentioned electronic device.
[0063] As Figure 4 shown, taking the mobile phone 300 as an example of the above-mentioned electronic device, the decoding method may include the following steps:
[0064] S401: The mobile phone 300 obtains a first data group at a first moment.
[0065] Among them, the first data group includes N data frames, and N is an integer greater than 1. For example, N can be equal to 2, 3, 4, etc., and is not limited herein.
[0066] Exemplarily, in one scenario, the above-mentioned N data frames can all be voice frames. The mobile phone 300 can extract voice frames from the voice signal in response to the detected voice signal. When the mobile phone 300 extracts N voice frames at the first moment, these N voice frames form the first data group.
[0067] Exemplarily, assume that N equals 3, and the mobile phone 300 can receive a voice signal input by the user in real time. When receiving the voice signal 1 at the 1st ms, the mobile phone 300 can extract the corresponding voice frame 1 from the voice signal 1; when receiving the voice signal 2 at the 2nd ms, the mobile phone 300 can extract the corresponding voice frame 2 from the voice signal 2; when receiving the voice signal 3 at the 3rd ms, the mobile phone 300 can extract the corresponding voice frame 3 from the voice signal 3. When the mobile phone 300 extracts the voice frames 1 - 3, the voice frames 1 - 3 can be used as the first data group, and the following step S402 can be continued. Based on this example, the above-mentioned first moment refers to the moment corresponding to the 3rd ms. That is to say, the mobile phone 300 can obtain a group of data for decoding processing in units of 3 voice frames.
[0068] In some embodiments, the above-mentioned voice signal can be a voice signal input by the user to the mobile phone 300 (for example, the voice signal input by the user during a speech in a square, or the voice signal input by the user during a lecture in a classroom). Or, the above-mentioned voice signal can also be a voice signal transmitted from another electronic device or server to the mobile phone 300, which is not limited herein.
[0069] It should be noted that each pronunciation in the above-mentioned voice signal can correspond to one voice frame or multiple voice frames. For example, when the voice signal is the pronunciation of "I" and "love", "I" can correspond to the voice frames 1 and 2 in the first data group, and "love" can correspond to the voice frame 3 in the first data group. For another example, when the voice signal is the pronunciation of "I" and "love", "I" can correspond to the voice frame 1 in the first data group, and "love" can correspond to the voice frames 2 and 3 in the first data group.
[0070] Exemplarily, in another scenario, the above-mentioned N data frames can also be image frames. The mobile phone 300 can obtain an image signal and extract image frames from the image signal. When the mobile phone 300 extracts N image frames at the first moment, the N image frames form the first data group.
[0071] Similarly, the process of the mobile phone 300 obtaining the first data group when the data frame is an image frame is the same as the process of the mobile phone 300 obtaining the first data group when the data frame is a voice frame, and will not be elaborated herein.
[0072] In some embodiments, the above-mentioned image signal can be an image signal input by the user using the mobile phone 300 to scan a target file (for example, the image signal input by the mobile phone 300 when scanning an article, or the image signal input by the mobile phone 300 when scanning a speech PPT). Or, the above-mentioned image signal can also be an image signal transmitted from another electronic device or server to the mobile phone 300, which is not limited herein.
[0073] S402: The mobile phone 300 performs two-way decoding on the first data group and outputs the first state parameter and the semantic vector corresponding to each of the N data frames, where the first state parameter is used to indicate the semantic features of the first data group.
[0074] Exemplarily, the first data group may include voice frame 1 and voice frame 2. When the mobile phone 300 obtains voice frame 1, it can process voice frame 1 into frame vector X1; when the mobile phone 300 obtains voice frame 2, it can process voice frame 2 into frame vector X2.
[0075] Exemplarily, the way for the mobile phone 300 to process voice frame 1 into frame vector X1 can be: first, convert the voice signal of voice frame X1 in the time domain into a voice signal in the frequency domain through fast Fourier transform, then filter the voice signal in the frequency domain, and then extract the voice feature vector (such as the timbre, pitch, and loudness of the voice) from the filtered voice signal in the frequency domain, so as to process voice frame 1 containing voice features into frame vector X1. In addition, the way for the mobile phone 300 to process voice frame 2 into frame vector X2 is the same as the way for the mobile phone 300 to process voice frame 1 into frame vector X1, which will not be elaborated here.
[0076] In another scenario, the first data group may include image frame 1 and image frame 2. When the mobile phone 300 obtains image frame 1, it can process image frame 1 into frame vector X1; when the mobile phone 300 obtains image frame 2, it can process image frame 2 into frame vector X2.
[0077] Exemplarily, the way for the mobile phone 300 to process image frame 1 into frame vector X1 can be: first, convert the image signal of image frame 1 in the time domain into an image signal in the frequency domain through fast Fourier transform, then remove the noise from the image signal in the frequency domain, and then extract the image feature vector (such as the color and brightness of the pixels) from the image signal after removing the noise, so as to process image frame 1 containing image features into frame vector X1. In addition, the way for the mobile phone 300 to process image frame 2 into frame vector X2 is the same as the way for the mobile phone 300 to process image frame 1 into frame vector X1, which will not be elaborated here.
[0078] After processing each voice frame or image frame in the first data group into the corresponding frame vector, the mobile phone 300 can perform two-way decoding on the obtained frame vectors, and then output the first state parameter and the semantic vector corresponding to each frame vector.
[0079] Next, taking frame vector X1 and frame vector X2 as examples, combined with Figure 5 the above two-way decoding process will be illustrated by examples.
[0080] When the data frames in the first data group are processed into frame vectors X1 and frame vector X2, the forward sequence corresponding to the first data group is "frame vector X1, frame vector X2", and the reverse sequence corresponding to the first data group is "frame vector X2, frame vector X1". That is to say, each frame vector arranged in the forward order according to the time sequence in the first data group is the forward sequence, and each frame vector arranged in the reverse order according to the time sequence in the first data group is the reverse sequence.
[0081] A first bidirectional decoder is provided in the mobile phone 300, and the mobile phone 300 can decode the frame vector X1 and the frame vector X2 in the first bidirectional decoder. As Figure 5 shown, the above-mentioned first bidirectional decoder includes a decoding unit B1, a decoding unit B2, a decoding unit F1, and a decoding unit F2. The decoding unit B1 is connected in series with the decoding unit B2, and the decoding unit F1 and the decoding unit F2 are connected in series.
[0082] Please continue to refer to Figure 5 , for example, the mobile phone 300 can input the frame vector X1 to the decoding unit F1 at t = 1, and the mobile phone 300 can also input the frame vector X2 to the decoding unit F2 at t = 2, that is, input the forward sequence of each data frame in the first data group.
[0083] After receiving the frame vector X1, the decoding unit F1 can decode X1 in the forward sequence "frame vector X1, frame vector X2" based on the pre-set first initial forward state parameter FS0, and output an intermediate state parameter FS1 and a forward decoding result y1. Among them, FS0 can be a vector of all zeros, and the intermediate state parameter FS1 represents the semantic feature corresponding to the frame vector X1 in the forward sequence "frame vector X1, frame vector X2".
[0084] For example, at the moment of t = 1, the decoding unit F1 can output the intermediate state parameter FS1 according to the first formula based on the pre-set first initial forward state parameter FS0, and the intermediate state parameter FS1 includes FC t and Fh t .
[0085] Among them, the first formula is: FC t = FC t-1 *f t +g t *i t , Fh t = o t *tanh(FC t ).
[0086] FC t-1 is the activation layer parameter of the first initial forward state parameter, f t is the forgetting parameter configured for the decoding unit F1, g tTransition parameters configured for decoding unit F1, i t Input parameters configured for decoding unit F1, o t Output parameters configured for decoding unit F1, FC t Activation layer parameters for intermediate state parameter FS1, Fh t Hidden layer parameters for intermediate state parameter FS1.
[0087] Among them, f t = σ(W Fxf x t + W Fhf Fh t-1 ), g t = tanh(W Fxg x t + W Fhg Fh t-1 ), o t = σ(W Fx0 x t + W Fh0 Fh t-1 ), where W Fxi , W Fhi , W Fxf , W Fhf , W Fxg , W Fhg , W FXO , W FHO are all set weight parameters, x t is the frame vector, Fh t-1 is the hidden layer parameter of the first initial forward state parameter,
[0088] It can be seen from the first formula that the decoding unit F1 can output the activation layer parameter FC t-1 of the intermediate state parameter FS1 according to the activation layer parameter FC t of the first initial forward state parameter; the decoding unit F1 can output the hidden layer parameter Fh t-1 of the intermediate state parameter FS1 according to the hidden layer parameter Fh t of the first initial forward state parameter.
[0089] After receiving the frame vector X2, the decoding unit F2 can decode X2 in the forward sequence "frame vector X1, frame vector X2" based on the intermediate state parameter FS1, and output the first forward state parameter FS2 and the forward decoding result y2. Among them, the first forward state parameter FS2 is used to indicate the semantic features of the forward sequence "frame vector X1, frame vector X2" and the semantic features of the frame vector "X2". It can be seen that the first forward state parameter characterizes the semantic features of the forward sequence corresponding to the first data group and the semantic features of the frame vector at the end of the forward sequence corresponding to the first data group.
[0090] Similarly, at the moment of t = 2, the manner in which the decoding unit F2 outputs the first forward state parameter FS2 based on the intermediate state parameter FS1 is the same as the above manner, and will not be elaborated here.
[0091] Similarly, still as Figure 5 shown, the reverse sequence of "frame vector X1, frame vector X2" is "frame vector X2, frame vector X1". While the first bidirectional decoder decodes the forward sequence "frame vector X1, frame vector X2" unidirectionally, it can also decode the reverse sequence "frame vector X2, frame vector X1" unidirectionally. For example, the mobile phone 300 can input the frame vector X2 to the decoding unit B1 at t = 1, and the mobile phone 300 can input the frame vector X1 to the decoding unit B2 at t = 2, that is, input the reverse sequences of each data frame in the first data group.
[0092] After receiving the frame vector X2, the decoding unit B1 can decode X2 in the reverse sequence "frame vector X2, frame vector X1" based on the preset first initial reverse state parameter BS0, and output the intermediate state parameter BS1 and the reverse decoding result y3. Among them, the first initial reverse state parameter BS0 can also be an all-zero vector, and the intermediate state parameter BS1 characterizes the semantic features corresponding to the frame vector X2 in the reverse sequence "frame vector X2, frame vector X1".
[0093] Similarly, at the moment of t = 1, the decoding unit B1 can output the intermediate state parameter BS1 according to the second formula based on the preset first initial forward state parameter BS0. Among them, the intermediate state parameter BS1 includes BC t and Bh t .
[0094] Among them, the second formula is: BC t = BC t-1 * f t + g t * i t , Bh t = o t * tanh(BC t ).
[0095] Among them, BCt-1 The activation layer parameter for the first initial reverse state parameter, f t The forgetting parameter configured for decoding unit B1, g t The transition parameter configured for decoding unit B1, i t The input parameter configured for decoding unit B1, o t The output parameter configured for decoding unit B1, BC t The activation layer parameter for the intermediate state parameter BS1, Bh t The hidden layer parameter for the intermediate state parameter BS1.
[0096] Among them, i t = σ(W Bxi x′ t + W Bhi Bh t-1 ), f t = σ(W Bxf x′ t + W Bhf Bh t-1 ), g t = tanh(W Bxg x′ t + W Bhg Bh t-1 ), o t = σ(W Bx0 x′ t + W Bh0 Bh t-1 ), where W Bxi , W Bhi , W Bxf , W Bhf , W Bxg , W Bhg , W BXO , W BHO are all set weight parameters, x′ t is the frame vector, Bh t-1 is the hidden layer parameter of the first initial reverse state parameter,
[0097] It can be seen from the second formula that decoding unit B1 can output the activation layer parameter BC t-1 of the intermediate state parameter BS1 according to the activation layer parameter BC t of the first initial reverse state parameter. Decoding unit B1 can output the hidden layer parameter Bh t-1 of the intermediate state parameter BS1 according to the hidden layer parameter Bh t of the first initial reverse state parameter.
[0098] After receiving the frame vector X1, the decoding unit B2 can decode X1 in the reverse sequence "frame vector X2, frame vector X1" based on the intermediate state parameter BS1, and output the first reverse state parameter BS2 and the reverse decoding result y4. The first reverse state parameter BS2 is used to indicate the semantic features of the reverse sequence "frame vector X2, frame vector X1" and the semantic features of the frame vector "X1" in the reverse sequence. It can be seen that the first reverse state parameter characterizes the semantic features of the reverse sequence corresponding to the first data group and the semantic features of the frame vector at the end of the reverse sequence corresponding to the first data group. It should be noted that the first forward state parameter and the first reverse state parameter constitute the above-mentioned first state parameter.
[0099] Similarly, at t = 2, the method for the decoding unit B2 to output the first reverse state parameter BS2 based on the intermediate state parameter BS1 is the same as the above method, and will not be elaborated here.
[0100] Finally, still as Figure 5 shown, the first bidirectional decoder can determine the semantic vector Y1 corresponding to the frame vector X1 according to the forward decoding result y1 output by the decoding unit F1 and the reverse decoding result y3 output by the decoding unit B1.
[0101] Exemplarily, the first bidirectional decoder can analyze the forward decoding result y1 and the reverse decoding result y3 according to the softmax function to determine the first probability of the forward decoding result y1 as Y1 and the second probability of the reverse decoding result y3 as Y1. If the first probability is greater than the second probability, the first bidirectional decoder outputs the forward decoding result y1 as Y1; if the first probability is less than the second probability, the first bidirectional decoder outputs the reverse decoding result y2 as Y1.
[0102] Similarly, still as Figure 5 shown, the first bidirectional decoder can determine the semantic vector Y2 corresponding to the frame vector X2 based on the forward decoding result y2 and the reverse decoding result y4.
[0103] Exemplarily, the first bidirectional decoder can analyze the forward decoding result y2 and the reverse decoding result y4 according to the softmax function to determine the third probability of the forward decoding result y2 as Y2 and the fourth probability of the reverse decoding result y4 as Y2. If the third probability is greater than the fourth probability, the first bidirectional decoder outputs the forward decoding result y2 as Y2; if the third probability is less than the fourth probability, the first bidirectional decoder outputs the reverse decoding result y4 as Y2.
[0104] Subsequently, the first bidirectional decoder can input the semantic vector Y1 and the semantic vector Y2 into a preset text generation model, so that the text generation model can output the text corresponding to the semantic vector Y1 and the semantic vector Y2 according to the semantic vector Y1 and the semantic vector Y2.
[0105] S403: The mobile phone 300 updates the first state parameter to a second state parameter, where the second state parameter discards the semantic features of the reverse sequence of the first data group.
[0106] Exemplarily, please continue to refer to Figure 5 , the first bidirectional decoder can output a first forward state parameter by performing the above step S402. The first forward state parameter includes a first hidden layer parameter h1 and a first activation layer parameter c1. Among them, the first hidden layer parameter h1 is used to indicate the semantic features of the data frame at the end of the forward sequence of the first data group, and the first activation layer parameter c1 is used to indicate the semantic features of the forward sequence of the first data group. The first reverse state parameter includes a second hidden layer parameter h2 and a second activation layer parameter c2. Among them, the second hidden layer parameter h2 is used to indicate the semantic features of the data frame at the end of the forward sequence of the first data group, and the second activation layer parameter c2 is used to indicate the semantic features of the forward sequence of the first data group.
[0107] Exemplarily, the first hidden layer parameter h1 can be a first matrix, and the second hidden layer parameter h2 can be a second matrix. Please continue to refer to Figure 5 , in one implementation, the first bidirectional decoder can splice the first hidden layer parameter h1 and the second hidden layer parameter h2 to obtain (h1h2), and use (h1h2) as the third hidden layer parameter of the second forward state parameter FS3; and, the first bidirectional decoder can use the first activation layer parameter C1 as the activation layer parameter of the second forward state parameter FS3. The first bidirectional decoder can also use the spliced (h1h2) as the third hidden layer parameter of the second reverse state parameter BS3; the first bidirectional decoder can also use the first activation layer parameter C1 as the activation layer parameter of the second reverse state parameter BS3.
[0108] In another implementation, the first bidirectional decoder can superimpose the first hidden layer parameter h1 and the second hidden layer parameter h2 to obtain (h1 + h2), and use (h1 + h2) as the third hidden layer parameter of the second forward state parameter FS3; similarly, the first bidirectional decoder can use the first activation layer parameter C1 as the activation layer parameter of the second forward state parameter FS3. The first bidirectional decoder can also use the spliced (h1 + h2) as the third hidden layer parameter of the second reverse state parameter BS3; the first bidirectional decoder can also use the first activation layer parameter C1 as the activation layer parameter of the second reverse state parameter BS3.
[0109] Understandably, when the second forward state parameter FS3 is [(h1h2), C1], the second reverse state parameter BS3 is also [(h1h2), C1]; or when the second forward state parameter FS3 is [(h1+h2), C1], the second reverse state parameter BS3 is also [(h1+h2), C1]. That is to say, the second forward state parameter FS3 is equal to the second reverse state parameter BS3.
[0110] Exemplarily, since the activation layer parameter in the second forward state parameter FS3 is equal to the first activation layer parameter of the first forward state parameter FS2. Therefore, the activation layer parameter in the second forward state parameter FS3 is also used to indicate: the semantic features of the forward sequence of the first data group. Since the third hidden layer parameter of the second forward state parameter FS3 is equal to the result of concatenating or superimposing the first hidden layer parameter h1 and the second hidden layer parameter h2. Thus, the hidden layer parameter in the second forward state parameter FS3 is used to indicate: the semantic features of the data frame at the end of the forward sequence of the first data group and the semantic features of the data frame at the end of the reverse sequence of the first data group. Similarly, since the second reverse state parameter BS3 is equal to the second forward state parameter FS3. Therefore, the meanings of the activation layer parameter and the hidden layer parameter in the second reverse state parameter BS3 are the same as those of the activation layer parameter and the hidden layer parameter in the second forward state parameter, which will not be elaborated here.
[0111] Understandably, the semantic features of the forward sequence of the first data group can express the semantics of the speech signal 1 input by the user at the first moment, and the semantics of the speech signal 1 generally has a high correlation with the semantics of the subsequent speech signal 2 input by the user. Then, when the mobile phone 300 subsequently obtains the second data group corresponding to the speech signal 2, the mobile phone 300 can retain the semantic features of the forward sequence of the first data group in the second state parameter. For example, retain the first activation layer parameter C1 of the first forward state parameter FS2 in the first state parameter in the second state parameter, so that the mobile phone 300 can use the second state parameter to decode the second data group more accurately.
[0112] The semantic features of the data frame at the end of the forward sequence of the first data group and the semantic features of the data frame at the end of the reverse sequence of the first data group are the reference basis for generating the semantic vector corresponding to the data frame at the end of the first data group. Generally, the semantics of the data frame at the end of the first data group has a high correlation with the data frame at the beginning of the second data group. Then, when the mobile phone 300 subsequently obtains the second data group, the mobile phone 300 can retain the semantic features of the data frame at the end of the forward sequence of the first data group and the semantic features of the data frame at the end of the reverse sequence of the first data group in the second state parameter. For example, the result after concatenating or stacking the first hidden layer parameter h1 of the first forward state parameter FS2 and the first hidden layer parameter h2 of the first reverse state parameter BS2 in the first state parameter is retained in the second state parameter, so that the mobile phone 300 can use the second state parameter to decode the second data group more accurately.
[0113] However, the semantic features of the reverse sequence of the first data group cannot express the semantics of the voice signal 1 input by the user at the first moment, and the semantic features of the reverse sequence of the first data group generally have no direct association with the semantics of the subsequent voice signal 2 input by the user. Then, when the mobile phone 300 subsequently obtains the second data group corresponding to the voice signal 2, the mobile phone 300 can discard the semantic features of the reverse sequence of the first data group, that is to say, the semantic features of the reverse sequence of the first data group will not be retained in the second state parameter, so that the mobile phone 300 can use the second state parameter to decode the second data group more accurately.
[0114] S404: The mobile phone 300 obtains the second data group at the second moment.
[0115] The second data group includes M data frames, the second moment is later than the first moment, and M is an integer greater than 1. For example, M can be equal to 2, 3, 4, etc., which is not limited here. Under the above conditions, M can be less than N, greater than N, or equal to N, which is not limited here. Among them, the way the mobile phone 300 obtains the second data group is the same as the way it obtains the first data group, which will not be elaborated here.
[0116] It should be noted that the category of the data frames in the first data group needs to be the same as the category of the data frames in the second data group. For example, when the data frames in the first data group are voice frames, the data frames in the second data group should also be voice frames; when the data frames in the second data group are image frames, the data frames in the second data group should also be image frames. In addition, there is no sequential order between S404 and S403.
[0117] Furthermore, the semantic features of the second data group can be associated with those of the first data group. For example, the semantic features of the first data group are the semantic features corresponding to the two Chinese characters "I" and "love"; the semantic features of the second data group are the semantic features corresponding to the two characters "China". It can be seen that the semantic features corresponding to "I" and "love" are associated with the semantic features corresponding to "China". For another example, the semantic features of the first data group are the semantic features corresponding to the two words "I" and "love"; the semantic features of the second data group are the semantic features corresponding to "china". It can be seen that the semantic features corresponding to "I" and "love" are associated with the semantic features corresponding to "china".
[0118] S405: The mobile phone 300 performs bidirectional decoding on the second data group according to the second state parameter, and outputs a semantic vector corresponding to each of the M data frames and a third state parameter. The third state parameter is used to indicate the semantic features of the second data group.
[0119] Exemplarily, as Figure 5 shown, it is assumed that a second bidirectional decoder is further provided in the mobile phone 300. The second bidirectional decoder includes a decoding unit F3, a decoding unit F4, a decoding unit B3, and a decoding unit B4. The decoding unit F3 and the decoding unit F4 are connected in series, and the decoding unit B3 and the decoding unit B4 are connected in series.
[0120] It is assumed that the second data group includes a data frame 3 and a data frame 4. The second bidirectional decoder can process the data frame 3 and the data frame 4 in the second data group into frame vectors 3 and 4 in the same way as the first bidirectional decoder processes the first data group into frame vectors.
[0121] The forward sequence of these two frame vectors is: "X3, X4". Furthermore, the second bidirectional decoder performs unidirectional decoding on the forward sequence "X3, X4". The difference is that during the decoding of the forward sequence "X1, X2", the initial forward state parameter is input to the decoding unit F1; while during the decoding of the forward sequence "X3, X4", the second forward state parameter is input to the decoding unit F2. The subsequent decoding method for the forward sequence "X3, X4" is the same as the above decoding method for the forward sequence "X1, X2".
[0122] For example, in the decoding unit F3, X3 in the forward sequence "X3, X4" can be decoded according to the second forward state parameter FS3, and the intermediate state parameter FS4 and the forward decoding result y5 are output. In the decoding unit F4, X4 in the forward sequence "X3, X4" is decoded based on the intermediate state parameter FS4, and the third forward state parameter FS5 and the forward decoding result y6 are output. Among them, the third forward state parameter FS5 is used to indicate the semantic features of the forward sequence "X3, X4" and the semantic features of "X4". It can be seen that the third forward state parameter is used to indicate the semantic features of the forward sequence corresponding to the second data group and the semantic features of the frame vector corresponding to the data frame at the end of the forward sequence corresponding to the second data group.
[0123] Please continue to refer to Figure 5 , the reverse sequence of the frame vector X3 and the frame vector X4 is: "X4, X3". The second bidirectional decoder can perform unidirectional decoding on the reverse sequence "X4, X3" while decoding the forward sequence "X3, X4". The difference is that during the decoding of the reverse sequence "X1, X2", the initial reverse state parameter is input to the decoding unit B1; while during the decoding of the reverse sequence "X4, X3", the second forward state parameter is input to the decoding unit B2. The subsequent decoding method for the reverse sequence "X4, X3" is the same as the above decoding method for the reverse sequence "X2, X1".
[0124] For example, in the decoding unit B3, X4 in the reverse sequence "X4, X3" can be decoded based on the second reverse state parameter BS3, and the intermediate state parameter BS4 and the reverse decoding result y7 are output. In the decoding unit B4, X3 in the reverse sequence "X4, X3" is decoded based on the intermediate state parameter BS4, and the third reverse state parameter BS5 and the reverse decoding result y8 are output. Among them, the third forward state parameter BS5 is used to indicate the semantic features of the reverse sequence "X4, X3" and the semantic features of "X3". It can be seen that the third reverse state parameter is used to indicate the semantic features of the reverse sequence corresponding to the second data group and the semantic features of the frame vector corresponding to the data frame at the end of the reverse sequence corresponding to the second data group. It should be noted that the third forward state parameter and the third reverse state parameter constitute the above-mentioned third state parameter.
[0125] Finally, the second bidirectional decoder can determine the semantic vector Y3 corresponding to the frame vector X3 according to the forward decoding result y5 and the reverse decoding result y7, and the second bidirectional decoder can determine the semantic vector Y4 corresponding to the frame vector X4 based on the forward decoding result y6 and the reverse decoding result y8.
[0126] Subsequently, the second bidirectional decoder can input the semantic vectors Y3 and Y4 into the text generation model, and the text generation model can output the text corresponding to the semantic vectors Y3 and Y4 according to the semantic vectors Y3 and Y4.
[0127] Based on the above example, it can be seen that the second bidirectional decoder decodes X3 in the forward sequence "X3, X4" according to the second forward state parameter FS3, and decodes X4 in the reverse sequence "X4, X3" according to the second backward state parameter BS3. The second forward state parameter FS3 and the second backward state parameter BS3 together constitute the second state parameter. Thus, when the second bidirectional decoder performs bidirectional decoding on the second data group, it refers to the second state parameter generated from the first state parameter, that is, refers to the semantic features of the first data group, so that the semantic vectors corresponding to the second data group and the semantic vectors corresponding to the first data group are coherent, and thus the accuracy of the finally output text is higher.
[0128] Based on the above, in the embodiment of the present application, bidirectional decoding can be performed on the first data group in the first bidirectional decoder. And bidirectional decoding can be performed on the second data group according to the first state parameter in the second bidirectional decoder.
[0129] It can be understood that when the second bidirectional decoder performs bidirectional decoding on the second data group at the second moment, the decoding of the first data group has been completed. Then, the first bidirectional decoder can obtain a new data group (such as the third data group) at the same moment (i.e., the second moment) and perform bidirectional decoding on the third data group. In this way, the mobile phone 300 can perform bidirectional decoding on multiple data groups simultaneously at the same moment, thereby improving the decoding efficiency.
[0130] Alternatively, the mobile phone 300 can also be provided with a bidirectional decoder, which can perform bidirectional decoding on the first data group obtained first and then perform bidirectional decoding on the obtained second data group. The embodiment of the present application does not make any restrictions on this.
[0131] As Figure 6 shown, the decoding method provided by the embodiment of the present application may further include:
[0132] S406: The mobile phone 300 updates the third state parameter to the fourth state parameter, and the fourth state parameter discards the semantic features of the reverse sequence of the second data group.
[0133] S407: The mobile phone 300 obtains the third data group at the third moment.
[0134] The above third data group includes L data frames, the third moment is later than the second moment, and L is an integer greater than 1. It can be understood that the manner of obtaining the third data group is the same as the manner of obtaining the first data group or the second data group described above, and will not be elaborated here.
[0135] The manner of updating the third state parameter to the fourth state parameter is the same as the manner and effect of updating the first state parameter to the second state parameter described above (reference can be made to Figure 5 ), which will not be elaborated here. There is no sequence precedence between S407 and S406 either.
[0136] S408: The mobile phone 300 performs bidirectional decoding on the third data group according to the fourth state parameter, and outputs a semantic vector corresponding to each of the L data frames.
[0137] Exemplarily, performing bidirectional decoding on the third data group according to the fourth state parameter includes: performing unidirectional decoding on the forward sequence of the third data group and performing unidirectional decoding on the reverse sequence of the third data group.
[0138] On the one hand, performing unidirectional decoding on the forward sequence of the third data group, the output result is: the third forward state parameter. The third forward state parameter is used to indicate the semantic features of the forward sequence of the third data group. The forward sequence of the third data group is used to indicate the frame sequence obtained by sorting the L data frames according to the timing relationship of the obtained L data frames.
[0139] On the other hand, performing unidirectional decoding on the reverse sequence of the third data group, the output result is: the third reverse state parameter. The third reverse state parameter is used to indicate the semantic features of the reverse sequence of the third data group. The reverse sequence of the third data group is used to indicate the frame sequence obtained by sorting the L data frames according to the reverse timing relationship of the obtained L data frames.
[0140] Similarly, the manner of performing unidirectional decoding on the forward sequence of the third data group is the same as the manner of performing unidirectional decoding on the forward sequence of the second data group described above, which will not be elaborated here. The manner of performing unidirectional decoding on the reverse sequence of the third data group is the same as the manner of performing unidirectional decoding on the reverse sequence of the second data group described above, which will not be elaborated here either.
[0141] Similarly, when performing bidirectional decoding on the third data group, the fourth state parameter generated from the third state parameter is referred to, that is, the semantic features of the second data group are referred to, so that the semantic vector groups corresponding to the second data group and the third data group obtained are coherent, thereby increasing the length of the text with determined semantic coherence and accuracy.
[0142] Exemplarily, Figure 7 is a schematic structural diagram of the decoding device 700 provided in the embodiment of the present application. Figure 7 The provided decoding device 700 can be used to execute Figure 4 , Figure 6The decoding method in. It should be noted that the decoding device 700 provided in the embodiments of the present application has the same basic principle and technical effects as those in the above embodiments. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above embodiments.
[0143] The device 700 includes: an acquisition unit 701, a first decoding unit 702, a state update unit 703, and a second decoding unit 704. Among them, the acquisition unit 701 is used to execute the above S401 and S404. The first decoding unit 702 is used to execute the above S402. The state update unit 703 is used to execute the above S403. The second decoding unit 704 is used to execute the above S405.
[0144] In addition, the device 700 may further include a third decoding unit 704. The acquisition unit 701 may also be used to execute the above S407. The state update unit 703 may also be used to execute the above S406. The third decoding unit 704 may be used to execute the above S408.
[0145] Exemplarily, Figure 8 is a schematic structural diagram of the electronic device 800 provided in the embodiments of the present application. The following combines Figure 8 Specific introductions will be made to the various components of the electronic device 800:
[0146] Among them, the processor 801 is the control center of the electronic device 800, which can be a single processor or a collective term for multiple processing elements. For example, the processor 801 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0147] Optionally, the processor 801 can execute various functions of the electronic device 800 by running or executing software programs stored in the memory 802 and calling data stored in the memory 802. For example, the processor 801 can execute S401 - S408 in the above embodiments of the present application, which is not limited herein.
[0148] In a specific implementation, as an embodiment, the processor 801 may include one or more CPUs, for example Figure 8The CPU0 and CPU1 shown in
[0149] In a specific implementation, as an embodiment, the electronic device 800 may also include multiple processors, such as Figure 2 the processors 801 and 804 shown in
[0150] Among them, the memory 802 is used to store the software program for executing the solution of this application, and is controlled by the processor 801 for execution. The specific implementation manner can refer to the above method embodiment and will not be elaborated here.
[0151] Optionally, the memory 802 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 802 may be integrated with the processor 801 or may exist independently and be coupled to the processor 801 through the interface circuit of the electronic device 800 ( Figure 8 not shown in
[0152] The transceiver 803 is used for communication with other electronic devices. For example, if the electronic device 800 is an electronic device, the transceiver 803 can be used for communication with a network device or with another electronic device. Another example is that if the electronic device 800 is a network device, the transceiver 803 can be used for communication with an electronic device or with another network device.
[0153] Optionally, the transceiver 803 may include a receiver and a transmitter ( Figure 8 not shown separately in
[0154] Optionally, the transceiver 803 may be integrated with the processor 801 or exist independently, and is coupled to the processor 801 through an interface circuit ( Figure 8 not shown) of the electronic device 800. The embodiments of the present application do not make specific limitations on this.
[0155] It should be noted that Figure 8 the structure of the electronic device 800 shown in does not constitute a limitation on the decoding device. The actual decoding device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0156] In addition, for the technical effects of the electronic device 800, reference may be made to the technical effects of the decoding method described in the foregoing method embodiments, and details are not described herein again.
[0157] The embodiments of the present application further provide a computer-readable storage medium, in which computer program code is stored. When the processor executes the computer program code, the electronic device executes the method in the foregoing embodiments.
[0158] The embodiments of the present application further provide a computer program product. When the computer program product runs on an electronic device, the electronic device is caused to execute the method in the foregoing embodiments.
[0159] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example for illustration. In actual applications, the above functions may be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. For the specific working processes of the systems, devices, and units described above, reference may be made to the corresponding processes in the foregoing method embodiments, and details are not described herein again.
[0160] In each of the embodiments of the present application, the functional units may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0161] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs.
[0162] As described above, the above is only the specific implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of the present application should be covered by the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application should be subject to the protection scope of the claims.
Claims
1. A decoding method, characterized in that, Including: Obtaining a first data group at a first moment, the first data group including N data frames, where N is an integer greater than 1; Bidirectionally decoding the first data group to output a first state parameter and semantic vectors corresponding to each of the N data frames, where the first state parameter is used to indicate the semantic features of the first data group, the first state parameter includes a first forward state parameter and a first reverse state parameter, the first forward state parameter is used to indicate the semantic features of the forward sequence of the first data group and the semantic features of the data frame at the end of the forward sequence of the first data group, and the first reverse state parameter is used to indicate the semantic features of the reverse sequence of the first data group and the semantic features of the data frame at the end of the reverse sequence of the first data group; Obtaining a second data group at a second moment, the second data group including M data frames, the second moment being later than the first moment, where M is an integer greater than 1; Bidirectionally decoding the second data group according to the first state parameter to output semantic vectors corresponding to each of the M data frames.
2. The method according to claim 1, characterized in that, The bidirectionally decoding the first data group to output a first state parameter includes: Unidirectionally decoding the forward sequence of the first data group to output the first forward state parameter, where the forward sequence of the first data group is used to indicate the frame sequence obtained by sorting the N data frames according to the temporal relationship of the obtained N data frames; Unidirectionally decoding the reverse sequence of the first data group to output the first reverse state parameter, where the reverse sequence of the first data group is used to indicate the frame sequence obtained by sorting the N data frames according to the reverse temporal relationship of the obtained N data frames.
3. The method according to claim 2, wherein The bidirectionally decoding the second data group according to the first state parameter includes: Updating the first state parameter to a second state parameter, where the second state parameter discards the semantic features of the reverse sequence of the first data group; Bidirectionally decoding the second data group according to the second state parameter.
4. The method according to claim 3, characterized in that, The second state parameter includes a second forward state parameter and a second reverse state parameter, and the second forward state parameter is equal to the second reverse state parameter; The activation layer parameter in the second forward state parameter is used to indicate: the semantic features of the forward sequence of the first data group; The hidden layer parameter in the second forward state parameter is used to indicate: the semantic features of the data frame at the end of the forward sequence of the first data group and the semantic features of the data frame at the end of the reverse sequence of the first data group.
5. The method according to claim 4, wherein The first hidden layer parameter in the first forward state parameter is a first matrix; the second hidden layer parameter in the first reverse state parameter is a second matrix; The hidden layer parameter in the second forward state parameter is a matrix formed by splicing or superimposing the first matrix and the second matrix.
6. The method according to claim 2, wherein, The first forward state parameter includes a first hidden layer parameter and a first activation layer parameter. The first hidden layer parameter is used to indicate the semantic features of the data frame at the end of the forward sequence of the first data group, and the first activation layer parameter is used to indicate the semantic features of the forward sequence of the first data group; The first reverse state parameter includes a second hidden layer parameter and a second activation layer parameter. The second hidden layer parameter is used to indicate the semantic features of the data frame at the end of the reverse sequence of the first data group, and the second activation layer parameter is used to indicate the semantic features of the reverse sequence of the first data group.
7. According to the method according to any one of claims 1-6, characterized in that, The method further includes: While outputting the semantic vectors corresponding to each of the M data frames, output a third state parameter, where the third state parameter is used to indicate the semantic features of the second data group; Obtain a third data group at a third moment. The third data group includes L data frames. The third moment is later than the second moment, and L is an integer greater than 1; Bidirectionally decode the third data group according to the third state parameter, and output the semantic vectors corresponding to each of the L data frames.
8. The method according to claim 7, characterized in that, The third state parameter includes a third forward state parameter and a third reverse state parameter. Among them, the bidirectional decoding of the third data group and outputting the third state parameter includes: Unidirectionally decode the forward sequence of the third data group to output the third forward state parameter. The third forward state parameter is used to indicate the semantic features of the forward sequence of the third data group and the semantic features of the data frame at the end of the forward sequence of the third data group. The forward sequence of the third data group is used to indicate the frame sequence obtained by sorting the L data frames according to the temporal relationship of the obtained L data frames; Unidirectionally decode the reverse sequence of the third data group to output the third reverse state parameter. The third reverse state parameter is used to indicate the semantic features of the reverse sequence of the third data group and the semantic features of the data frame at the end of the reverse sequence of the third data group. The reverse sequence of the third data group is used to indicate the frame sequence obtained by sorting the L data frames according to the reverse temporal relationship of the obtained L data frames.
9. The method according to claim 8, wherein The bidirectional decoding of the third data group according to the third state parameter includes: Update the third state parameter to a fourth state parameter, and the fourth state parameter discards the semantic features of the reverse sequence of the second data group; Bidirectionally decode the third data group according to the fourth state parameter.
10. The method according to claim 1, wherein The bidirectional decoding of the first data group includes: in a first decoder, bidirectionally decode the first data group; The bidirectional decoding of the second data group according to the first state parameter includes: in a second decoder, bidirectionally decode the second data group according to the first state parameter.
11. The method according to claim 1, characterized in that, All of the N data frames are speech frames. The obtaining of the first data group at the first moment includes: In response to the input speech signal, extract speech frames from the speech signal; When N speech frames are extracted at the first moment, the N speech frames form the first data group; Alternatively, when the N data frames are all image frames, obtaining the first data group at the first moment includes: In response to an input image signal, extracting an image frame from the image signal; When N image frames are extracted at the first moment, the N image frames form the first data group.
12. The method according to claim 1, wherein The semantic feature of the first data group is associated with the semantic feature of the second data group.
13. An electronic device, characterized in that, Including: One or more processors; A memory; And one or more computer programs, wherein the one or more computer programs are stored on the memory, and when the computer programs are executed by the one or more processors, the electronic device executes the decoding method according to any one of claims 1-12.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program or instruction, and when the computer program or instruction runs on a computer, the computer executes the decoding method according to any one of claims 1-12.
15. A computer program product, characterized in that, The computer program product includes: a computer program or instruction, and when the computer program or instruction runs on a computer, the computer executes the decoding method according to any one of claims 1-12.
Citation Information
Patent Citations
Video encoding method and device, video decoding method and device, and electronic equipment
CN111526370A
Bidirectional prediction method and image decoding device
WO2020138958A1