Reporting and data fields
Leverage audio and analytics data in reports.
Medallia Speech data feeds directly into Experience Cloud reports and populates a range of data fields for in-depth analysis. This guide details the data fields populated by Speech and the reports impacted by Speech transcripts and audio.
Speech audio and analytics data in reports
Responses Form reports and Responses Feed modules can show a media player with the transcript for playback of the audio file associated with the transcribed call. To enable the media player, turn on the Display media from collected survey property, and include the Medallia Speech transcript (q_speech_comment) field in the report or module configuration to embed the transcribed call.
Responses Form reports and Responses Feed modules automatically include topics and sentiments identified by Text Analytics.
For example, this image shows data for a voice call in a Responses Feed module:
The bar above the transcript shows sentiment for phrases spoken by the customer and the agent, as well as events such as call silence and overtalk (when the agent and customer speak at the same time).
The media bar enables you to hear the actual call, while the transcription pane shows the text transcript between the agent and the customer, displayed in a common chat format. As you play the call, the text transcript scrolls so you can read along as you listen. If you want to listen to and read different parts of the call, either click a new place in the media bar or click a piece of the text transcript.
The Topics & Sentiments pane shows the topics and sentiments captured by Text Analytics for this call, with the number on the right indicating the amount of matching phrases. For example, in the call above, the Agent Quality - Politeness topic was mentioned a total of 7 times. Since that topic is expanded, we can see that each mention of the topic was by the agent. The module shows a timestamp button for each mention. Click a timestamp to open the call at that point in both the media bar and the transcript pane.
To edit the topics and sentiment captured for the call, click the pencil icon to the right of the media bar.
With Speech, Text Analytics report filters can include a Speaker dropdown, which allows you to filter the report to show All comments, only comments from the Agent, or only comments from the Customer. If you also add the Media Type filter, users can filter responses to show only responses with audio feedback. For more information, see Filters and control panels.
Speech call data fields
This section lists the required and optional call data fields, the optional call metadata fields, and the A-fields used in reporting.
Required call data fields
This table lists the root-level fields for Speech. These fields are required to be imported for each call record.
| Field | Description |
|---|---|
| Call identifier | A unique record identifier. It can be the Universal Call ID (UCID) |
| Speech file name | Name of the audio file associated with record. Note: AWS S3 supports the use of forward slashes in file names to simulate folders. If your company uses this feature, you must include the full path in the Speech File Name. For example, audio/2020-07-03!1000/T15996_A.wav. |
| Unit ID | ID of the agent that handled the call (typically the last agent the customer is transferred to if there are multiple agents). This must match the ID that is included in the organizational hierarchy for the agent. Note: While this field is always required, you can supplement with additional unit fields through the transfer of custom metadata. For example, if your company is using Apps (formerly known as Best Practice Packages) that have a different unit field, you can include that field as metadata. |
| Response date | The Interaction date and time in this format: yyyy-MM-dd HH:mm:ssZZ (for example, 2016-01-01 11:30:00-0800) |
| Speech vendor | The speech-to-text transcription engine to use for the call. |
Optional call data fields
This table lists the optional Event and System fields that can be imported for each call record. Medallia recommends that you include these additional metadata values to ensure your audio is being transcribed more precisely to meet your business needs.
| Field | Description |
|---|---|
| Call recording URL | URL of the call interaction recording. |
| Vertical | The Speech vertical model to use, such as Call Center. |
| Survey language | The language locale spoken by the customer during the call, such as en_US. |
| Agent survey language | The language locale spoken by the agent during the call, such as en_US. This field is available only when the Vendor is Engine1. |
| Apply diarization | Determines whether diarization will be applied to the audio file during processing. This field applies to mono files only, which must be diarized. |
| Agent channel | Determines which of the two channels (0 or 1) is associated with the agent. The other channel is associated with the customer. Note: The initiator of a call is assigned to channel 0. For inbound calls, set the Agent Channel to 1. For outbound calls, set the Agent Chanel to 0. |
| Substitutions | Substitutions which must be done in the file. |
| Redaction | Determines whether redaction is performed on the audio and its transcription. |
| First name | First name of the customer. |
| Last name | Last name of the customer. |
| Email address of the customer. | |
| Phone | Phone number of the customer. |
Call metadata fields
This table lists the call metadata Event and Feedback fields used by Speech to provide helpful information about each call record. These fields are automatically populated by Medallia as part of transcribing audio files.
| Field | Description |
|---|---|
| Medallia product | The Medallia product associated with the record, populated by Medallia as Speech. |
| Signal type | The type of feedback, populated by Medallia as Call Transcripts / Speech Analytics. |
| Channel | The channel associated with the feedback, populated by Medallia as Voice Call. |
| Signal source | The source of the feedback signal, populated by Medallia as Medallia. |
| Medallia Speech transcript | The text transcript of the voice call, in VTT format. |
Scores and related fields
Speech uses the following fields to record information about voice data and populate reports. These fields are derived from the JSON transcript of call audio.
If you provide audio or video for Speech to process, we create the transcript and derive the field values from there.
If you provide a completed JSON transcript instead of audio, you must ensure the required values are present in the transcript in order for Speech to map or calculate the field values. For information on formatting a JSON transcript and the fields required, see ASR transcript schema.
For more information on the Experience Cloud fields listed below, see System fields.
- Agent acoustic emotion trend
- The change detected in the agent's acoustic emotion from the beginning to the end of the call.
Possible values include Improving, Similar, Worsening, or Too short if the call is shorter than 90 seconds.
Note: The first 45 seconds of the call are not considered for this calculation because the beginning of a call tends to be limited to standard greetings; omitting those greetings more accurately describes acoustic emotion. - Agent average talk streak
- The average time for all streaks of agent speech, measured in seconds.
A streak is an uninterrupted chain of spoken language by a single speaker. A streak begins with the first word of a call, or with the next word spoken after a pause of three seconds or more. A streak ends with a pause of three seconds or more, or with the end of the call.
- Agent clarity
- A number from 0-100 representing the ASR engine's confidence that it correctly transcribed the audio channel with agent speech.
Higher numbers indicate better clarity, which is calculated from a confidence score based on the ASR engine's own estimate of its transcription quality. Clarity is not accuracy, which must be determined using human-generated transcripts, and clarity is not an indicator of the speaker's elocution.
- Agent longest talk streak
- The total time for the longest streak of agent speech, measured in seconds.
A streak is an uninterrupted chain of spoken language by a single speaker. A streak begins with the first word of a call, or with the next word spoken after a pause of three seconds or more. A streak ends with a pause of three seconds or more, or with the end of the call.
- Agent non-silence ratio
- Agent talk time during the call as a percentage of overall call duration.
- Agent overtalk incidents
- Number of overtalk incidents by the agent during a call.
- Agent overtalk ratio
- Agent overtalk time during the call as a percentage of overall call duration.
- Agent talk rate
- The average rate of agent speech for the entire call, expressed as words per minute. Customer speech and significant pauses in the conversation are not included in this calculation.
- Agent talk rate trend
- The change detected in the agent's talk rate from the beginning to the end of the call. Expressed as a number between 0 and 1.0, where 1.0 represents no change in talk rate and .3 represents a 70% decrease in talk rate.
Calculated by comparing the agent's talk rate during the first third of the call to the agent's talk rate during the last third of the call.
- Agent vs Customer talk rate
- The ratio of agent talk rate to customer talk rate.
- Customer acoustic emotion trend
- The change detected in the customer's acoustic emotion from the beginning to the end of the call.
Possible values include Improving, Similar, Worsening, or Too short if the call is shorter than 90 seconds.
Note: The first 45 seconds of the call are not considered for this calculation because the beginning of a call tends to be limited to standard greetings; omitting those greetings more accurately describes acoustic emotion. - Customer average talk streak
- The average time for all streaks of customer speech, expressed in seconds.
A streak is an uninterrupted chain of spoken language by a single speaker. A streak begins with the first word of a call, or with the next word spoken after a pause of three seconds or more. A streak ends with a pause of three seconds or more, or with the end of the call.
- Customer clarity
- A number from 0-100 representing the ASR engine's confidence that it correctly transcribed the audio channel with customer speech.
Higher numbers indicate better clarity, which is calculated from a confidence score based on the ASR engine's own estimate of its transcription quality. Clarity is not accuracy, which must be determined using human-generated transcripts, and clarity is not an indicator of the speaker's elocution.
- Customer longest talk streak
- The total time for the longest streak of customer speech, expressed in seconds.
A streak is an uninterrupted chain of spoken language by a single speaker. A streak begins with the first word of a call, or with the next word spoken after a pause of three seconds or more. A streak ends with a pause of three seconds or more, or with the end of the call.
- Customer non-silence ratio
- Customer talk time during the call as a percentage of total call duration.
- Customer overtalk incidents
- The count of overtalk incidents by the customer during a call.
- Customer overtalk ratio
- Customer overtalk time during the call as a percentage of total call duration.
- Customer talk rate
- The average rate of customer speech for the entire call, expressed as words per minute. Agent speech and significant pauses in the conversation are not included in this calculation.
- Customer talk rate trend
- The change detected in the customer's talk rate from the beginning to the end of the call. Expressed as a number between 0 and 1.0, where 1.0 represents no change in talk rate and .3 represents a 70% decrease in talk rate.
Calculated by comparing the customer's talk rate during the first third of the call to the customer's talk rate during the last third of the call.
- Luhn presence
- When redaction is turned on for Speech transcripts. Whether or not the Luhn algorithm (which identifies credit card numbers) was applied to the transcript.
- Media duration
- The duration in seconds of the call.
- Overtalk
- The percentage of the call during which the customer and agent spoke simultaneously.
- Overall agent acoustic emotion
- The agent's average acoustic emotion during the call.
Options include Strongly positive, Positive, Neutral, Negative, and Strongly negative, or Too short for calls shorter than 90 seconds.
Note: The first 45 seconds of the call are not considered for this calculation because the beginning of a call tends to be limited to standard greetings; omitting those greetings more accurately describes acoustic emotion. - Overall customer acoustic emotion
- The customer's average acoustic emotion during the call.
Options include Strongly positive, Positive, Neutral, Negative, and Strongly negative, or Too short for calls shorter than 90 seconds.
Note: The first 45 seconds of the call are not considered for this calculation because the beginning of a call tends to be limited to standard greetings; omitting those greetings more accurately describes acoustic emotion. - Silence (incidents)
- A count of the times during a call when no speech was detected.
- Silence (percentage)
- The percentage of the call during which no speech was detected.
