The Ultimate Guide to BIN Sponsorship

Home Business Magazine Online

With BIN sponsorships being a quick alternative for companies to access major card schemes and build scalable payment products, whilst also allowing for flexibility, it has continued to gain traction within the payments industry and is increasingly becoming the preferred payment route for FinTechs.

We have put together this guide to provide clarity on BIN sponsorship and help you gain a better understanding of what BIN sponsorship is, the different routes that can be taken within a BIN sponsorship, and the benefits BIN sponsorship could have for your company.

What Is BIN Sponsorship?

A BIN sponsorship is a single compliant connection between a card scheme network and an organisation, which allows for transactions to be processed and cardholder funds to be settled. Having a BIN sponsor, such as Transact Payments, enables businesses to perform many operations of a financial institution through them, without having to obtain banking licenses.

We’ve broken down the term BIN sponsorship:

BIN – stands for Banking Identification Number, and it refers to the initial set of four to six numbers within the long card number that is on the front of a payment card. These numbers are used to identify the banking institution that has issued the card and are used to match transactions to the issuer to ensure a smooth customer journey.

Sponsorship – refers to the relationship between the client and a sponsor (the issuer of the BIN). To be a BIN sponsor, the issuer has to be a direct scheme member of the major card schemes (i.e. MasterCard and Visa). As scheme members, banking institutions can sponsor their clients to use scheme-approved BINs and account ranges.

When it comes to issuing cards through a BIN sponsorship, you must abide by the local requirements and regulations in the country a business is wishing to issue a card in, and the BIN sponsor must have licenses to issue in those countries. Despite what some companies may claim, there are currently no global issuers of BIN sponsorships, due to these requirements. At Transact Payments, we provide active BIN and Sub-BIN ranges in 25 European countries and 17 different currencies.

To obtain a payment card, other parties are needed, and as a BIN sponsor, we have the connections to ensure all relevant parties are involved.

When Are BIN Sponsorships Used?

There are several situations, where a BIN sponsorship may be sought after by a business. These situations include when a business is looking to add:

  • Corporate expense cards
  • General spend cards
  • Gift cards
  • Money transfer cards
  • Travel and foreign exchange cards
  • Payroll and payroll-plus cards
  • Affiliate commission cards

Depending on the needs of the business, and the intended use, these may be issued as physical, virtual prepaid, debit, credit or tokenized cards.

What Routes Can Be Taken During BIN Sponsorship?

When it comes to BIN sponsorship, there are a number of routes that can be taken, these include a dedicated BIN sponsorship route, settlement only route and service provider route.

What Are the Benefits of BIN Sponsorships?

BIN Sponsorships allow businesses to access major card schemes without having to obtain their own licenses, which is typically a long and strenuous process. Instead, BIN sponsorships allow businesses to quickly access payment and major card schemes, by using their sponsor’s licenses.

Going through a BIN sponsor is beneficial, as they are able to guide a business on the laws and regulations of issuing payment cards in those countries and ensure that their programme adheres to the Scheme rules and regulatory requirements.

BIN Sponsors can also guide businesses through the BIN setup process and perform ongoing support and management to help you deliver a compelling programme.

Why Choose Transact Payments? 

You are in safe hands with the Transact Payments team, who have extensive knowledge of both regulatory and scheme requirements.

Transact Payments are licenced UK and European e-money institutions, and Principal members of both MasterCard and Visa. As a BIN sponsor, we can share our licenses with our clients, allowing them to perform operations of a financial institution without having to obtain banking licenses themselves.

We provide secure, flexible and innovative card and payment solutions. Our BIN sponsorship service is fully integrated with many leading processors, and we have the capabilities to integrate with any business’s proprietary processing system to enable a flexible, cost-efficient and fast program implementation, with ongoing management that can be relied upon.

Our solutions are used by regulated financial institutions, financial services companies, and marketing organizations.

We are transparent about everything we do at Transact Payments, and we are actively working to raise industry standards. When it comes to BIN sponsorship, we work with our clients to decide the best route, and if we feel BIN sponsorship is not worthwhile for your business, then we will be up front about it. Get in contact with our experts today to find out more about our BIN Sponsorship scheme.

The post The Ultimate Guide to BIN Sponsorship appeared first on Home Business Magazine.

Original source: https://homebusinessmag.com/businesses/success-tips/ultimate-guide-bin-sponsorship/

How to make money ironing

So, you want to spend your spare time ironing?

Let’s just say, better you than us!

Since that’s what most people think, there’s an extremely high demand for efficient ironing services that do quality work. People will rather pay extra to outsource their laundry than iron it themselves and that’s where you come in.

If you can get your hands on some coat hangers and know your way round an ironing board, this may well be the money-maker you’ve been looking for.

Read MoneyMagpie’s quick guide to starting your own ironing business below:

 

What’s involved in an ironing business?

Electric iron next to pile of folded laundry

In order to offer a high quality ironing service, you will need the essentials: a good iron and ironing board. It’s worth investing in good quality tools – it’ll make the labour easier and there’ll be less of a chance they damage your customers’ clothes. You should also have experience ironing, even if it’s just your own garments, and know the basics, for example, different heat settings for various types of fabrics. You’ll probably want to offer your clients higher ironing standards than you do for your own clothes.

There may already be an ironing service in your vicinity with plenty of customers, so delivering the best possible results is of the utmost importance. Do some market research to check what others offer in terms of pricing, services and delivery options. Then have a think about how you can top their offer.

Starting out ironing for cash

You basically have two options:

  1. Sign up with an agency,
  2. Or start your own ironing business.
ironing agencies

With the first option, it’s as simple as finding the closest agency and convincing them that you’re good at ironing. You can find it online through a quick search for iron agency and the name of your area.

Different agencies provide different services – some include alterations and repairs – so find one that best suits your skills.

When you sign up with an agency, they deal with administrative aspects such as insurance, however, if you must be registered as self-employed (and the agency should tell you this), you will have to consider some income tax issues which you can read about in this article on paying tax on your extra income.

Some agencies simply put you in touch with a local customer and leave you to arrange payments yourself.

Others have clients who typically need housekeeping, cleaning and ironing work. So, they match an ironer to their client’s needs and allow the two parties to agree a price.

If you decide to  collaborate with an agency, the work is normally part-time and often flexible.

Payment ranges from hourly rates to a fee per item or a fee per pound.

In general, you can expect to earn between:

  • £8 and £12 per hour
  • 50p and £1 per item
  • 50p and £1 per pound

Although some agencies won’t require you to drive, having wheels definitely makes your job easier. If you don’t already have a vehicle, read our article on car leasing.

starting your own business

The big advantage of running your own operation is that you get to keep all of the money! You also get to work whenever you want and pick your clients.

The downside is that you have to do all the work of advertising, invoicing, dealing with customers and so on.

However, if this doesn’t put you off, read on to see how you can get all of this going.

 

How do I start my own ironing service?

Woman ironing in the living room/lounge

Research

Start by scouring your local paper and in-shop windows to find out how stiff the competition is. Have a quick search online too and check any local Facebook sites – people often advertise on community or neighbourhood pages.

If there’re lots of ironing services available already, there probably won’t be room for any more.

If, however, you don’t notice any, great start. There may actually be demand for your services. A good way to get your business off the ground is to start small and gradually build it as you gain more clients.

For tax purposes, it’s important that you register as self-employed within three months of working for yourself. Visit Gov.uk to find out more.

Services

Start by offering a simple ironing service and then take things from there. You can also offer seasonal services, such as special offers for school uniforms in September or polishing up people’s outfits for the wedding season.

And you can branch out. If you’re a confident sewer, for instance, your talents will definitely come in handy. You can offer mending and even washing services too.

But it’s not just about getting stuff ironed. Remember – presentation is everything. It’s common practice to return a customer’s ironing in a clear plastic bag or on hangers, so be sure to have a constant supply of these. You can ask your clients to include their own hangers when you pick up their order but make sure you have some extras on hand.

A major service that most customers appreciate, is collection and delivery. It’s acceptable to charge a small fee for this, based on the cost of fuel and time spent driving.

On the plus side, when you register as self-employed, you can claim the tax back for the costs incurred when driving to and from the customer – so keep all your petrol receipts.

Costings

The starting-out cost of an ironing gig is low but setting up a budget is an essential part of any business – and it shouldn’t take too long.

All you have to do is draw up a list of expenses – including supplies, fuel, advertising, rent etc – and work out their total. When you quote customers, ensure that you cover all these bases and also add an amount for labour. It’s a good idea to compare your prices with other businesses offering ironing services. Many people in the ironing business admit they started out charging way too little and found it hard to increase their prices later as it posed the risk of losing good clients. Make sure you’re confident your fees reflect the work you put into your business.

Advertising

There’s lots of different ways in which you can advertise your ironing business but word of mouth remains the best. So, let all your friends, family and work colleagues know that you do this and ask them to help you spread the word.

It will also help to set up a Facebook page for your business. You  can use it to post regular updates and it’s a place where your customers can share their reviews. If they’re positive, it’ll encourage others to use your services.

You can also consider advertising in shop windows, libraries, your local paper, the Yellow Pages and, eventually, even setting up your own website.

 

Top tips for making money with ironing

Folded clothes on ironing board

Running an ironing business doesn’t have to be complicated but there are a few things you should keep in mind:

  • Get some business cards printed cheaply with Vistaprint and hand them out at every opportunity. You could also design some simple leaflets and deliver them to houses in your neighbourhood.
  • Smokers need not apply. If your customers get a whiff of cigarette smoke on their clothes you won’t be expecting any repeat business. It’s also a good idea to work away from strong cooking smells and pets (allergies).
  • Ironing in the summer can be hot work so pick up a cheap fan to keep you cool while you’re working.
  • Good professional indemnity insurance is a must in case you iron a hole in anyone’s designer suit. Accidents happen!
  • Advertise your services on free local sites like Facebook groups, local supermarket notice boards and cheaper places like Gumtree.
  • Websites such as Indeed or Totaljobs also post ironing jobs every so often so make sure you set up job alerts. Freelance sites like PeoplePerHour could also be a good way to find new clients.

Hopefully you now know everything you need to start ironing to ramp up your income. The beauty of this money making strategy is how easy and flexible it is to set up. Your ironing gig can be as small as only doing a few shirts for someone you know or it could be a proper full-time job. You can iron when kids are in bed or while watching TV. Happy ironing!

The post How to make money ironing appeared first on MoneyMagpie.

Original source: https://www.moneymagpie.com/make-money/make-money-ironing

How to Encourage Remote Employees to Move Closer to the Home Office

Home Business Magazine Online

In the digital age, the world of work is continually evolving, allowing for increasingly flexible working arrangements. Remote work, a concept once unimaginable, is now a reality for many employees. Also, companies often have multiple offices across the nation and even globally, but the home office often houses vital operations and decision-making processes.

Thus, it inevitably attracts the need for the best talent. However, persuading employees to relocate and move closer to the home office is often challenging. Not everyone is readily willing to uproot their lives for professional reasons.

This guide provides insights into how companies might incentivize remote employees to relocate, making the home office an attractive destination.

Offering Relocation Assistance

Securing new accommodation, finding suitable schools for their kids, managing the sale and purchase of homes, and orchestrating a physical move can be daunting and a deterrent to moving.

Acknowledging these challenges and offering relocation assistance could become a game-changer in the decision-making process for your employees. Consider providing support in securing suitable new accommodation, perhaps even offering temporary lodging while employees settle in, and providing resources for school scouting to alleviate some of the immediate stress associated with moving.

For most individuals, one of the key concerns with relocation is the financial burden. To address this, your company could offer to cover moving costs. You can use moving cost calculators to calculate your moving cost and get a fair estimate of how much it will cost you in every specific case.

Providing Opportunities for Growth

It’s crucial to remember that just as you have aspirations for your company’s growth, your employees have career progression goals. You can leverage this mutual inclination towards progress to incentivize their relocation to the home office.

For instance, you could offer a qualified employee a promotion to a managerial position from a lower-ranking role. Given the inherent prestige and increased responsibilities associated with higher positions, this prospect could serve as a powerful motivation for relocation.

In scenarios where immediate promotions may not be feasible, highlighting the potential for future advancement at the home office could be equally effective.

Implementing Flexible Work Arrangements

In some instances, complete relocation may not be feasible for everyone. This reality calls for implementing flexible work arrangements, which could serve as an attractive compromise. For instance, you may arrange for employees to work from the home office a few days per week, allowing them to retain the benefits of remote work during the remaining days.

This approach may particularly apply to employees whose families cannot relocate entirely. By tailoring arrangements according to each employee’s flexibility, you optimize their output without disregarding their personal circumstances.

It’s a win-win scenario that fosters productivity, engagement, and employee satisfaction. However, it’s crucial to remember that this approach requires effective communication, understanding, and mutual agreement.

Recognizing and Rewarding the Sacrifice

Recognizing and rewarding the courage and commitment of employees who choose to relocate can help foster satisfaction and loyalty. For example, publicly celebrating these employees as committed and dedicated individuals during company meetings or on internal communication platforms could boost morale and reinforce positive behavior.

Monetary rewards also wield substantial persuasive power. So, consider offering bonuses or pay hikes to employees post-relocation while ensuring that the increased pay matches the cost of living in the new location.

Research indicates that most employees may consider leaving an employer for better pay; likewise, they may be more inclined to relocate if the financial benefits are attractive.

Wrapping Up

The pursuit of employee relocation hinges on acknowledging the challenges your employees face during the relocation and offering tangible solutions to simplify the process.

Whether through financial assistance, career growth opportunities, flexible work arrangements, or rewards and recognition, every strategic move should be centered on making the relocation attractive enough for your employees to consider.

The post How to Encourage Remote Employees to Move Closer to the Home Office appeared first on Home Business Magazine.

Original source: https://homebusinessmag.com/businesses/how-to-guides-businesses/how-to-encourage-remote-employees-move-closer-home-office/

Are Hyperbaric Chambers Safe for Homes?

Home Business Magazine Online

Are you considering buying a hyperbaric chamber for use in your home? It’s an investment that is not to be taken lightly and it’s important to consider the pros, cons, safety, and efficacy of these pressurized chambers. Hyperbaric therapy has been increasing in popularity over recent years as celebrities have adopted the health trend; but how safe are these chambers and can they improve your overall well-being?

What Is a Hyperbaric Oxygen Chamber?

A hyperbaric oxygen chamber is a medical protocol that involves breathing pure oxygen in a pressurized room or tube. The high pressure of the chamber allows the body to absorb more oxygen than it would at normal atmospheric pressure. This increased oxygen saturation can be beneficial for a variety of medical conditions.

The protocol typically takes place in a specialized chamber that can be pressurized to several times the normal atmospheric pressure. The patient lies on a bed inside the chamber, and a mask or hood is placed over the nose and mouth to deliver pure oxygen. The patient breathes in the oxygen-rich air, and the increased pressure forces more oxygen into the bloodstream.

Hyperbaric oxygen therapy (HBOT) is often used to treat conditions that involve a lack of oxygen to the tissues, such as carbon monoxide poisoning, burns, and certain types of wounds that are slow to heal.

How Can You Use HBOT at Home?

Hyperbaric oxygen therapy (HBOT) is typically administered by trained medical professionals in a hospital or clinical setting. It is not generally recommended to use HBOT at home due to the specialized equipment and expertise required to safely administer HBOT. However, with Oxyhelp chambers, you can now enjoy the benefits of HBOT at home.

Sports 

HBOT at home can be used to help athletes of all levels improve their overall performance. By increasing the oxygen concentration in the blood, HBOT can increase energy levels and endurance, reduce recovery times from injuries and workouts, and boost mental alertness. It is especially effective for those involved in high-intensity or repetitive sports such as running and cycling, where long training sessions can take a toll on the body.

The increased oxygen also helps to reduce inflammation, swelling, and injuries caused by overuse or intense physical activity. By using HBOT at home, athletes can maximize their efforts and ensure that they stay at the top of their game.

Beauty and Wellness

HBOT has long been recognized for its therapeutic and regenerative properties, which is why it has grown in popularity among the beauty and wellness communities. The use of HBOT at home can help to reduce wrinkles, improve skin tone and texture, as well as stimulate collagen production. It is also known to have a positive effect on overall wellness by helping to reduce inflammation, boost immunity, alleviate pain and improve sleep quality.

Additionally, HBOT has been found to help detoxify the body by allowing oxygen molecules to effectively penetrate deep into the cells, which can help support healthy organ function. All of these benefits make using an HBOT at home a great addition to any beauty and wellness routine.

Anti-Aging 

Home hyperbaric oxygen therapy (HBOT) can be used as an anti-aging protocol. HBOT involves breathing in pure oxygen at higher-than-normal atmospheric pressure, which increases the amount of oxygen available to bodily tissues and helps to promote healing and growth. This increased oxygen level is believed to help skin cells repair damage from free radicals, reduce wrinkles and improve skin tone.

In addition to reducing the visible signs of aging, HBOT can help detoxify the body, stimulate collagen production and reduce inflammation. It is also believed to help increase energy levels and cognitive function, as well as promote overall health and wellness.

What Makes Oxyhelp Chambers Best for Homes?

Safer Oxygen Concentration

Oxyhelp Hyperbaric Chambers are one of the best choices for home use due to their safety features. They provide a high concentration of oxygen while still being safe, as they only use 95% oxygen instead of 100%. This lowered amount helps to reduce the risk of combustion, making them safer than other hyperbaric chambers.

Additionally, Oxyhelp is designed for easy use, with adjustable pressure and a range of settings. This makes it perfect for anyone who wants to experience the health benefits of hyperbaric oxygen therapy in their own home.

Oxyhelp
Photo Credit: Oxyhelp

Separate Concentrators

Oxyhelp hyperbaric chambers are the best for homes because of their safety features. They have an advanced air monitoring system that constantly measures oxygen levels and ensures optimal flow. In addition, they feature a patented low-noise design to ensure a peaceful environment during protocol.

The use of medical-grade tubing also guarantees the highest quality of oxygen delivery for optimal results. Furthermore, the chambers are easy to use and maintain, making them ideal for home use. All of these features make Oxyhelp hyperbaric chambers the best option for home use.

They provide a safe and comfortable environment while delivering medical-grade oxygen to help improve health and well-being. With our advanced safety features, ease of use, and medical-grade oxygen delivery, Oxyhelp Hyperbaric Chambers are the perfect choice for home use.

Easy Sliding Doors

Oxyhelp Hyperbaric Chambers are the best choice for home use due to their exceptional safety features. The chamber will depressurize in 40 seconds, guaranteeing a quick escape in an emergency. Furthermore, Oxyhelp chambers have easy sliding doors that make getting in and out of them easier than ever.

Oxyhelp
Photo Credit: Oxyhelp

Additionally, they are designed with multiple safety features, such as oxygen monitors to ensure the user is never exposed to toxic gas levels. Finally, Oxyhelp chambers are constructed with durable materials that can withstand heavy use and prolonged periods of operation.

Easy-to-use Interface

Oxyhelp hyperbaric chambers have been designed with convenience in mind. They are easy to set up and use, allowing you to take control of your oxygen therapy with the simple touch of a button. A convenient user interface also allows you to easily adjust settings such as the pressure level and protocol duration.

Additionally, our chambers can be controlled with your smartphone, which means you can adjust settings without having to leave the comfort of your own home. Therefore, Oxyhelp hyperbaric chambers are the ideal choice for anyone looking for an effective and convenient oxygen therapy solution.

If you’re looking for a way to improve your health and wellness at home, hyperbaric oxygen therapy is a great option. There are many benefits to HBOT, including improved sports performance, better skin health, and anti-aging effects. Oxyhelp chambers are the best choice for HBOT at home because they offer a safer oxygen concentration, separate concentrators, easy sliding doors, and an easy-to-use interface. If you’re interested in trying HBOT at home, Oxyhelp is the perfect solution.

The post Are Hyperbaric Chambers Safe for Homes? appeared first on Home Business Magazine.

Original source: https://homebusinessmag.com/lifestyles/health-and-fitness/hyperbaric-chambers-safe-homes/

Google’s 2023 Search quality rater guidelines update: Here’s what changed

It’s been nearly a year since Google last updated its Search Quality Rater guidelines. 

Unlike previous edits to the Search Quality Guidelines, which have introduced significant, new concepts (like the new E for Experience last year), the latest updates to the Search Quality Guidelines seem much more focused on user intent and needs met. 

Google is:

  • Refining what it means to provide high-quality search results.
  • Helping quality raters understand why certain results are more helpful than others. 

This level of nuance can help explain why we see certain volatility during core updates (as well as periods outside of announced algorithm updates).

If search quality raters have given Google ample evidence that its results are not meeting user expectations, this can lead to substantial intent shifts during core updates. 

Looking at search results for the same query, before and after major Google updates, makes a lot more sense when you understand the granularity with which Google approaches understanding the intent behind a query, and what it means to have high-quality, helpful content

Here are high-level insights into what has changed in the latest search quality rater guidelines update.

More guidance around rating page quality for forums & Q&A pages

Google added a new block of text to instruct raters on how to rate the quality of forum and Q&A pages, specifically in situations where the discussions are either brand new, or drifting into “combative,” “misleading” or “spammy content.” 

A forum page defaults to “medium” if it’s simply a new page that hasn’t had time to collect answers. But older posts without answers should be rated as low quality. 

Google mentions “decorum” a few times in this section, indicating that combative discussions that show a lack of respect should be rated as low quality. 

  • This is a good reminder that the quality of comments on a given page can impact the overall quality of that page’s content, assuming Google can crawl and index the content in the comments. Often, comment sections are neglected or unmoderated, and if they become problematic, insulting or disrespectful, this can negatively impact an otherwise good-quality page.
Alphabet Inc.
(Page 76)

Additionally, the new version of the Search Quality Guidelines includes a visual example of what a “medium” quality forum page looks like on Reddit. The question is only 9 hours old and has no answers, so it defaults to a score of “medium.”

It’s worth noting that Google implies there are no other low-quality characteristics on this page that otherwise could lean the score towards “low quality” for new discussions. 

APPLE
(Page 77)

Below is the example Reddit URL Google used in this case: 

Digital marketing

A quick addition about the importance of a location to a query

Google added a short snippet about the importance of user location to understanding a query. For searches looking for nearby places, location is important, whereas generic questions like “how does gravity work” have the same answer, regardless of the user’s location. 

Food Retail & Distribution (NEC)
(Page 84)
  • Interestingly, Google felt this nuance was worth adding to the Rater Guidelines. While it seems self-explanatory, it’s true that the extent to which a search query has a built in “local intent” can have a major impact on the types of results that would best answer that query. 

Expanding on “Minor Interpretations”

Google added a deeper explanation about its definitions for “Minor interpretations.”

Minor interpretations describe a situation when a query can have multiple meanings, and the minor interpretations are the least likely to be the commonly expected meanings of the query.

Within minor interpretations, Google introduced:

  • “Reasonable minor interpretations,” which help “fewer users” but are still helpful for search results. 
  • “Unlikely minor interpretations,” which are theoretically possible but highly unlikely. 

“No chance interpretations” are interpretations of a query that are incredibly unlikely for the user to be looking for. Google provides the example of an “overheated pet” when the searcher types “hot dog” (although I feel that this interpretation is more plausible than a “no chance” rating!).

Google
(Page 87)

Google also added some new visual examples of how to interpret these definitions. For example, an “unlikely minor interpretation” of the search query “Apple” would be the U.S. city, Apple, Oklahoma.

Google Search
(Page 88)

Further defining ‘Know Simple’ and ‘Do’ queries

Google added several new examples of what types of queries are not “Know Simple” queries.

Know Simple” queries are defined as queries that seek a very specific answer, like a fact or a diagram, that can be answered in a small amount of space, like one or two sentences. 

Google added three new examples of queries that are not Know Simple queries: when users want to browse or explore a topic, find inspiration related to a topic, or are seeking personal opinions and perspectives from real people.

Harvard
(Page 90)

What makes this addition interesting: The language used here is quite similar to the language Google uses when describing the value of SGE (Search Generative Experience). For example, the following language comes from the main SGE page:

  • “Dive deeper on a topic in a conversational way.”
  • “Access the high-quality results and perspectives that you expect from Google.”
  • “I want to know what people think to help me make a decision”

Perhaps – and this is purely speculation – the feedback Google gets from quality raters about whether queries can be classified as “Know Simple” (or not) can help them understand when to trigger SGE.

Along the same lines, Google added two new examples of “Know Simple Queries” and “Know Queries” – the bottom two rows of the below table – to provide additional context about when a query is easily answered or when the answer is more open ended.

Home Business
(Page 91)

Google also added more examples of “Do” queries to the table below, starting with [shape of you video] and all queries below it. The three new examples represent queries that would be best answered with videos, images or how-to guides.

internet marketing
(Page 91)

Further refining user intent

Google introduced new language around user intent by classifying what is an unlikely user intent for a set of keywords. 

In the table below, Google added a second column to explain unlikely intents for the keywords Harvard and Walmart. This limits the intent of these keywords to a more reasonable user intent, rather than a completely open-ended set of possible answers.

Users searching for “Harvard” could be looking for various details about the university, but are probably not looking for a specific course. 

internet search
(Page 95)

Examples of Google’s SERP features that highly meet user needs

Google provides a table with various examples of search queries, the user’s location, and the user’s intent. It then shows the search result, and rates the extent to which the result met the expectations of the user (“Needs met”).

Google also offers an explanation about why these particular results are ranked as “highly meeting” the expectations of users. 

In this new version of the QRG, Google added and adjusted some of the examples in the “Highly Meets (HM)” results, the highest possible rating of meeting user needs “for most queries.” 

One example of a new example Google added to this list is one where the user is looking for “nearby coffee shops.” Google provides a screenshot of a Google Maps local pack with three coffee shops listed, and explains why this result highly, but not fully, meets the user’s expectation (it doesn’t list every possible coffee shop). 

Internet search engines
(Page 114)

Google even added a TikTok video as an example of a result that highly meets the needs of users looking an “around the world tutorial” for soccer. 

Local search

These are just some of the new examples, which seem to provide a more modern view of different results, both from external sites as well as Google’s own SERP features.

Dig deeper. An SEO guide to understanding E-E-A-T

The post Google’s 2023 Search quality rater guidelines update: Here’s what changed appeared first on Search Engine Land.

Original source: https://searchengineland.com/google-2023-search-quality-rater-guidelines-update-changes-434981

Quality rater and algorithmic evaluation systems: Are major changes coming?

Crowd-sourced human quality raters have been the mainstay of the algorithmic evaluation process for search engines for decades. Still, a potential sea-change in research and production implementation could be on the horizon. 

Recent groundbreaking research by Bing (with some purported commercial implementation already) and a sharp uptick in closely related information retrieval research by others, indicates some big shake-ups are coming.

These shake-ups may have far-reaching consequences for both the armies of quality raters and potentially the frequency of algorithmic updates we see go live, too. 

The importance of evaluation

In addition to crawling, indexing, ranking and result serving for search engines is the important process of evaluation. 

How well does a current or proposed search result set or experimental design align with the notoriously subjective notion of relevance to a given query, at a given time, for a given search engine user’s contextual information needs?

Since we know relevance and intent for many queries are always changing, and how users prefer to consume information evolves, search result pages also need to change to meet both the searcher’s intent and preferred user interface. 

Some changes have predictable, temporal and periodic query intent shifts. For example, in the period approaching Black Friday, many queries typically considered informational might take sweeping commercial intent shifts. Similarly, a transport query like [Liverpool Manchester] might shift to a sports query on local match derby days. 

In these instances, an ever-expanding legacy of historical data supports a high probability of what users consider more meaningful results, albeit temporarily. These levels of confidence likely make seasonal or other predictable periodic results and temporary UI design shifting relatively straightforward adjustments for search engines to implement.

However, when it comes to broader notions of evolving “relevance” and “quality,” and for the purposes of experimental design changes too, search engines must know a proposed change in rankings after development by search engineers is truly better and more precise to information needs, than the present results generated. 

Evaluation is an important stage in search results evolution and vital to providing confidence in proposed changes – and substantial data for any adjustments (algorithmic tuning) to the proposed “systems,” if required. 

Evaluation is where humans “enter the loop” (offline and online) to provide feedback in various ways before roll-outs to production environments.

This is not to say evaluation is not a continuous part of production search. It is. However, an ongoing judgment of existing results and user activity will likely evaluate how well an implemented change continues to fare in production against an acceptable relevance (or satisfaction) based metric range. A metric range based on the initial human judge-submitted relevance evaluations.

In a 2022 paper titled, “The crowd is made of people: Observations from large-scale crowd labelling,” Thomas et al., who are researchers from Bing, allude to the ongoing use of such metric ranges in a production environment when referencing a monitored component of web search “evaluated in part by RBP-based scores, calculated daily over tens of thousands of judge-submitted labels.” (RBP stands for Rank-Biased Precision).

Human-in-the-loop (HITL)

Data labels and labeling

An important point before we continue. I will mention labels and labeling a lot throughout this piece, and a clarification about what is meant by labels and labeling will make the rest of this article much easier to understand:

I will provide you with a couple of real-world examples most people will be familiar with for breadth of audience understanding before continuing:

  • Have you ever checked a Gmail account and marked something as spam?
  • Have you ever marked a film on Netflix as “Not for me,” “I like this,” or “love this”?

All of these submitted actions by you create data labels used by search engines or in information retrieval systems. Yes, even Netflix has a huge foundation in information retrieval and a great information retrieval research team tool. (Note that Netflix is both information retrieval with a strong subset of that field, called “recommender systems.”)

By marking “Not for me” on a Netflix film, you submitted a data label. You became a data labeler to help the “system” understand more about what you like (and also what people similar to you like) and to help Netflix train and tune their recommender systems further.

Data labels are all around us. Labels markup data so it can be transformed into mathematical forms for measurement at scale. 

Enormous amounts of these labels and “labeling” in the information retrieval and machine learning space are used as training data for machine learning. 

“This image has been labeled as a cat.” 

“This image has been labeled as a dog… cat… dog… dog… dog… cat,” and so on. 

All of the labels help machines learn what a dog or a cat looks like with enough examples of images marked as cats or dogs.

Labeling is not new; it’s been around for centuries, since the first classification of items took place. A label was assigned when something was marked as being in a “subset” or “set of things.” 

Anything “classified” has effectively had a label attached to it, and the person who marked the item as belonging to that particular classification is considered the labeler.

But moving forward to recent times, probably the best-known data labeling example is that of reCAPTCHA. Every time we select the little squares on the image grid, we add labels, and we are labelers. 

We, as humans, “enter the loop” and provide feedback and data.

With that explanation out of the way, let us move on to the different ways data labels and feedback are acquired, and in particular, feedback for “relevance” to queries to tune algorithms or evaluate experimental design by search engines.

Implicit and explicit evaluation feedback

While Google refers to their evaluation systems in documents meant for the non-technical audience overall as “rigorous testing,” human-in-the-loop evaluations in information retrieval widely happen through implicit or explicit feedback.

Implicit feedback

With implicit feedback, the user isn’t actively aware they provide feedback. The many live search traffic experiments (i.e., tests in the wild) search engines carry out on tiny segments of real users (as small as 0.1%), and subsequent analysis of click data, user scrolling, dwell time and result skipping, fall into the category of implicit feedback. 

In addition to live experiments, the ongoing general click, scroll and browse behavior of real search engine users can also constitute implicit feedback and likely feed into “Learning to Rank (LTR) machine learning” click models. 

This, in turn, feeds into rationales for proposed algorithmic relevance changes, as non-temporal searcher behavior shifts and world changes lead to unseen queries and new meanings for queries. 

There is the age-old SEO debate around whether rankings change immediately before further evaluation from implicit click data. I will not cover that here other than to say there is considerable awareness of the huge bias and noise that comes with raw click data in the information retrieval research space and the huge challenges in its continuous use in live environments. Hence, the many pieces of research work around proposed click models for unbiased learning to rank and learning to rank with bias.

Regardless, it is no secret overall in information retrieval how important click data is for evaluation purposes. There are countless papers and even IR books co-authored by Google research team members, such as “Click Models for Web Search” (Chuklin and De Rijke, 2022). 

Google also openly states in their “rigorous testing” article:

“We look at a very long list of metrics, such as what people click on, how many queries were done, whether queries were abandoned, how long it took for people to click on a result and so on.”

And so a cycle continues. Detected change needed from Learning to Rank, click model application, engineering, evaluation, detected change needed, click model application, engineering, evaluation, and so forth.

Explicit feedback

In contrast to implicit feedback from unaware search engine users (in live experiments or in general use), explicit feedback is derived from actively aware participants or relevance labelers. 

The purpose of this relevance data collection is to mathematically roll it up and adjust overall proposed systems.

A gold standard of relevance labeling – considered near to a ground truth (i.e., the reality of the real world) of intent to query matching – is ultimately sought. 

There are various ways in which a gold standard of relevance labeling is gathered. However, a silver standard (less precise than gold but more widely available data) is often acquired (and accepted) and likely used to assist in further tuning.

Explicit feedback takes four main formats. Each has its advantages and disadvantages, largely about relevance labeling quality (compared with gold standard or ground truth) and how scalable the approach is.

Real users in feedback sessions with user feedback teams

Search engine user research teams and real users provided with different contexts in different countries collaborate in user feedback sessions to provide relevance data labels for queries and their intents. 

This format likely provides near to a gold standard of relevance. However, the method is not scalable due to its time-consuming nature, and the number of participants could never be anywhere near representative of the wider search population at large.

True subject matter experts / topic experts / professional annotators

True subject matter experts and professional relevance assessors provide relevance for query mappings annotated to their intents in data labeling, including many nuanced cases. 

Since these are the authors of the query to intent mappings, they know the exact intent, and this type of labeling is likely considered near to a gold standard. However, this method, similar to the user feedback research teams format, is not scalable due to the sparsity of relevance labels and, again, the time-consuming nature of this process. 

This method was more widely used before introducing the more scalable approach of crowd-sourced human quality raters (to follow) in recent times.

Search engines simply ask real users whether something is relevant or helpful

Real search engine users are actively asked whether a search result is helpful (or relevant) by search engines and consciously provide explicit binary feedback in the form of yes or no responses with recent “thumbs up” design changes spotted in the wild.

rustybrick on X - Google search result poll

Crowd-sourced human quality raters

The main source of explicit feedback comes from “the crowd.” Major search engines have huge numbers of crowd-sourced human quality raters provided with some training and handbooks and hired through external contractors working remotely worldwide. 

Google alone has a purported 16,000 such quality raters. These crowd-sourced relevance labelers and the programs they are part of are referred to differently by each search engine. 

Google refers to its participants as “quality raters” in the Quality Raters Program, with the third-party contractor referring to Google’s web search relevance program as “Project Yukon.” 

Bing refers to their participants as simply “judges” in the Human Relevance System (HRS), with third-party contractors referring to Bing’s project as simply “Web Content Assessor.” 

Despite these differences, participants’ purposes are primarily the same. The role of the crowd-sourced human quality rater is to provide synthetic relevance labels emulating search engine users across the world as part of explicit algorithmic feedback. Feedback often takes the form of a side-by-side (pairwise) comparison of proposed changes versus either existing systems or alongside other proposed system changes. 

Since much of this is considered offline evaluation, it isn’t always live search results that are being compared but also images of results. And it isn’t always a pairwise comparison, either. 

These are just some of the many different types of tasks that human quality raters carry out for evaluation, and data labeling, via third-party contractors. The relevance judges likely continuously monitor after the proposed change roll-out to production search, too. (For example, as the aforementioned Bing research paper alludes to.)

Whatever the method of feedback acquisition, human-in-the-loop relevance evaluations (either implicit or explicit) play a significant role before the many algorithmic updates (Google launched over 4,700 changes in 2022 alone, for example), including the now increasingly frequent broad core updates, which ultimately appear to be an overall evaluation of fundamental relevance revisited.


Get the daily newsletter search marketers rely on.

<input type="hidden" name="utmMedium" value="“>
<input type="hidden" name="utmCampaign" value="“>
<input type="hidden" name="utmSource" value="“>
<input type="hidden" name="utmContent" value="“>
<input type="hidden" name="pageLink" value="“>
<input type="hidden" name="ipAddress" value="“>

Processing…Please wait.

function getCookie(cname) {
let name = cname + “=”;
let decodedCookie = decodeURIComponent(document.cookie);
let ca = decodedCookie.split(‘;’);
for(let i = 0; i <ca.length; i++) {
let c = ca[i];
while (c.charAt(0) == ' ') {
c = c.substring(1);
}
if (c.indexOf(name) == 0) {
return c.substring(name.length, c.length);
}
}
return "";
}
document.getElementById('munchkinCookieInline').value = getCookie('_mkto_trk');


Relevance labeling at a query level and a system level

Despite the blog posts we have seen alerting us to the scary prospect of human quality raters visiting our site via referral traffic analysis, naturally, in systems built for scale, individual results of quality rater evaluations at a page level, or even at an individual rater level have no significance on their own. 

Human quality raters do not judge websites or webpages in isolation 

Evaluation is a measurement of systems, not web pages – with “systems” meaning the algorithms generating the proposed changes. All of the relevance labels (i.e., “relevant,” “not relevant,” “highly relevant”) provided by labelers roll up to a system level. 

“We use responses from raters to evaluate changes, but they don’t directly impact how our search results are ranked.”

– “How our Quality Raters make Search results better,” Google Search Help

In other words, while relevance labeling doesn’t directly impact rankings, aggregated data labeling does provide a means to take an overall (average) measurement of how well a proposed algorithmic change (system) might be, more precisely relevant (when ranked), with lots of reliance on various types of algorithmic averages.

Query-level scores are combined to determine system-level scores. Data from relevance labels is turned into numerical values and then into “average” precision metrics to “tune” the proposed system further before any roll-out to search engine users more broadly. 

How far from the expected average precision metrics engineers hoped to achieve with the proposed change is the reality when ‘humans enter the loop’?

While we cannot be entirely sure of the metrics used on aggregated data labels when everything is turned into numerical values for relevance measurement, there are universally recognized information retrieval ranking evaluation metrics in many research papers. 

Most authors of such papers are search engine engineers, academics, or both. Production follows research in the information retrieval field, of which all web search is a part.

Such metrics are order-aware evaluation metrics (where the ranked order of relevance matters, and weighting, or “punishing” of the evaluation if the ranked-order is incorrect). These metrics include:

  • Mean reciprocal rank (MRR).
  • Rank-biased precision (RBP).
  • Mean average precision (MAP).
  • Normalized and un-normalized discounted cumulative gain (NDCG and DCG respectively).

In a 2022 research paper co-authored by a Google research engineer, NDCG and AP (average precision) are referred to as a norm in the evaluation of pairwise ranking results:

“A fundamental step in the offline evaluation of search and recommendation systems is to determine whether a ranking from one system tends to be better than the ranking of a second system. This often involves, given item-level relevance judgments, distilling each ranking into a scalar evaluation metric, such as average precision (AP) or normalized discounted cumulative gain (NDCG). We can then say that one system is preferred to another if its metric values tend to be higher.”

– “Offline Retrieval Evaluation Without Evaluation Metrics,” Diaz and Ferraro, 2022

Information on DCG, NDCG, MAP, MRR and their commonality of use in web search evaluation and ranking tuning is widely available.

Victor Lavrenko, a former assistant professor at the University of Edinburgh, also describes one of the more common evaluation metrics, mean average precision, well:

“Mean Average Precision (MAP) is the standard single-number measure for comparing search algorithms. Average precision (AP) is the average of … precision values at all ranks where relevant documents are found. AP values are then averaged over a large set of queries…”

So it’s literally all about the averages judges submit from the curated data labels distilled into a consumable numerical metric versus the predicted averages hoped for after engineering and then tuning the ranking algorithms further.

Quality raters are simply relevance labelers

Quality raters are simply relevance labelers, classifying and feeding a huge pipeline of data, rolled up and turned into numerical scores for:

  • Aggregation on whether a proposed change is near an acceptable average level of relevance precision or user satisfaction.
  • Or determining whether the proposed change needs further tuning (or total abandonment).

The sparsity of relevance labeling causes a bottleneck

Regardless of the evaluation metrics used, the initial data is the most important part of the process (the relevance labels) since, without labels, no measurement via evaluation can take place.

A ranking algorithm or proposed change is all very well, but unless “humans enter the loop” and determine whether it is relevant in evaluation, the change likely won’t happen.

For the past couple of decades, in information retrieval widely, the main pipeline of this HITL-labeled relevance data has come from crowd-sourced human quality raters, which replaced the use of the professional (but fewer in numbers) expert annotators as search engines (and their need for speedy iteration) grew. 

Feeding yays and nays in turn converted into numbers and averages in order to tune search systems.

But scale (and the need for more and more relevance labeled data) is increasingly problematic, and not just for search engines (even despite these armies of human quality raters). 

The scalability and sparsity issue of data labeling presents a global bottleneck and the classic “demand outstrips supply” challenge.

Widespread demand for data labeling has grown phenomenally due to the explosion in machine learning in many industries and markets. Everyone needs lots and lots of data labeling. 

Recent research by consulting firm Grand View Research illustrates the huge growth in market demand, reporting:

“The global data collection and labeling market size was valued at $2.22 billion in 2022 and it is expected to expand at a compound annual growth rate of 28.9% from 2023 to 2030, with the market then expected to be worth $13.7 billion.”

This is very problematic. Particularly in increasingly competitive arenas such as AI-driven generative search with the effective training of large language models requiring huge amounts of labeling and annotations of many types.

Authors at Deepmind, in a 2022 paper, state:

 “We find current large language models are significantly undertrained, a consequence of the recent focus on scaling language models while keeping the amount of training data constant. …we find for compute-optimal training …for every doubling of model size the number of training tokens should also be doubled.” 

– “Training Compute-Optimal Large Language Models,” Hoffman et al. 

When the amount of labels needed grows quicker than the crowd can reliably produce them, a bottleneck in scalability for relevance and quality via rapid evaluation on production roll-outs can occur. 

Lack of scalability and sparsity do not fit well with speedy iterative progress

Lack of scalability was an issue when search engines moved away from the industry norm of professional, expert annotators and toward the crowd-sourced human quality raters providing relevance labels, and scale and data sparsity is once again a major issue with the status quo of using the crowd. 

Some problems with crowd-sourced human quality raters

In addition to the lack of scale, other issues come with using the crowd. Some of these relate to human nature, human error, ethical considerations and reputational concerns.

While relevance remains largely subjective, crowd-sourced human quality raters are provided with, and tested on, lengthy handbooks, in order to determine relevance. 

Google’s publicly available Quality Raters Guide is over 160 pages long, and Bing’s Human Relevance Guidelines is “reported to be over 70 pages long,” per Thomas et al.

Bing is much more coy with their relevance training handbooks. Still, if you root around, as I did when researching this piece, you can find some of the documentation with incredible detail on what relevance means (in this instance for local search), which looks like one of their judging guidelines in the depths online.

Efforts are made in this training to instill a mindset appreciative of the evaluator’s role as a “pseudo” search engine user in their natural locale. 

The synthetic user mindset needs to consider many factors when emulating real users with different information needs and expectations. 

These needs and expectations depend on several factors beyond simply their locale, including age, race, religion, gender, personal opinion and political affiliation. 

The crowd is made of people

Unsurprisingly, humans are not without their failings as relevance data labelers.

Human error needs no explanation at all and bias on the web is a known concern, not just for search engines but more generally in search, machine learning, and AI overall. Hence, the dedicated “responsible AI” field emerges in part to deal with combatting baked-in biases in machine learning and algorithms. 

However, findings in the 2022 large-scale study by Thomas et al., Bing researchers, highlight factors leading to reduced precision relevance labeling going beyond simple human error and traditional conscious or unconscious bias.

Even despite the training and handbooks, Bing’s findings, derived from “hundreds of millions of labels, collected from hundreds of thousands of workers as a routine part of search engine development,” underscore some of the less obvious factors, more akin to physiological and cognitive factors and contributing to a reduction in precision quality in relevance labeling tasks, and can be summarised as follows:

  • Task-switching: Corresponded directly with a decline in quality of relevance labeling, which was significant as only 28% of participants worked on a single task in a session with all others moving between tasks. 
  • Left side bias: In a side-by-side comparison, a result displayed on the left side was more likely to be selected as relevant when compared with results on the right side. Since pair-wise analysis by search engines is widespread, this is concerning.
  • Anchoring: Played a part in relevance labeling choices, whereby the relevance label assigned on the first result by a labeler is also much more likely to be the relevance label assigned for the second result. This same label selection appeared to have a descending probability of selection in the first 10 evaluated queries in a session. After 10 evaluated queries, the researchers found that the anchoring issue seemed to disappear. In this instance the labeler hooks (anchors) onto the first choice they make and since they have no real notion of relevance or context at that time, the probability of them choosing the same relevance label with the next option is high. This phenomenon disappears as the labeler gathers more information from subsequent pairwise sets to consider.
  • General fatigue of crowd-workers played a part in reduced precision labeling.
  • General disagreement between judges on which one of a pairwise result was relevant from the two options. Simply differing opinions and perhaps a lack of true understanding of the context of the intended search engine user.
  • Time of day and day of week when labeling was carried out by evaluators also plays a role. The researchers noted some related findings which appeared to correlate with spikes in reduced relevance labeling accuracy when regional celebrations were underway, and might have easily been considered simple human error, or noise, if not explored more fully.

The crowd is not perfect at all.

A dark side of the data labeling industry

Then there is the other side of the use of human crowd-sourced labelers, which concerns society as a whole. That of low-paid “ghost workers” in emerging economies employed to label data for search engines and others in the tech and AI industry.

Major online publications increasingly draw attention to this issue with headlines like:

And, we have Google’s own third-party quality raters protesting for higher pay as recently as February 2023, with claims of “poverty wages and no benefits.”

Add together all of this with the potential for human error, bias, scalability concerns with the status quo, the subjectivity of “relevance,” the lack of true searcher context at the time of query and the inability to truly determine whether a query has a navigational intent.

And we have not even touched upon the potential minefield of regulations and privacy concerns around implicit feedback.

How to deal with lack of scale and “human issues”?

Enter large language models (LLMs), ChatGPT and increasing use of machine-generated synthetic data.

Is the time right to look at replacing ‘the crowd’?

A 2022 research piece from “Frontiers of Information Access Experimentation for Research and Education” involving several respected information retrieval researchers explores the feasibility of replacing the crowd, illustrating the conversation is well underway.

Clarke et al. state: 

“The recent availability of LLMs has opened the possibility to use them to automatically generate relevance assessments in the form of preference judgements. While the idea of automatically generated judgements has been looked at before, new-generation LLMs drive us to re-ask the question of whether human assessors are still necessary.”

However, when considering the current situation, Clarke et al. raise specific concerns around a possible degradation in the quality of relevance labeling in exchange for huge scale potentials:

Concerns about reduced quality in exchange for scale?

“It is a concern that machine-annotated assessments might degrade the quality, while dramatically increasing the number of annotations available.” 

The researchers draw parallels between the previous major shift in the information retrieval space away from professional annotators some years before to “the crowd,” continuing:

“Nevertheless, a similar change in terms of data collection paradigm was observed with the increased use of crowd assessor…such annotation tasks were delegated to crowd workers, with a substantial decrease in terms of quality of the annotation, compensated by a huge increase in annotated data.”

They surmise that the feasibility of “over time” a spectrum of balanced machine and human collaboration, or a hybrid approach to relevance labeling for evaluations, may be a way forward. 

A wide range of options from 0% machine and 100% human right across to 100% machine and 0% human is explored.

The researchers consider options whereby the human is at the beginning of the workflow providing more detailed query annotations to assist the machine in relevance evaluation, or at the end of the process to check the annotations provided by the machines.

In this paper, the researchers draw attention to the unknown risks that may emerge through the use of LLMs in relevance annotation over human crowd usage, but do concede at some point, there will likely be an industry move toward the replacement of human annotators in favor of LLMs:

“It is yet to be understood what the risks associated with such technology are: it is likely that in the next few years, we will assist in a substantial increase in the usage of LLMs to replace human annotators.”

Things move fast in the world of LLMs

But much progress can take place in a year, and despite these concerns, other researchers are already rolling with the idea of using machines as relevance labelers.

Despite the concerns raised in the Clarke et al. paper around reduced annotation quality should a large-scale move toward machine usage occur, in less than a year, there has been a significant development that impacts production search.

Very recently, Mark Sanderson, a well-respected and established information retrieval researcher, shared a slide from a presentation by Paul Thomas, one of four Bing research engineers presenting their work on the implementation of GPT-4 as relevance labelers rather than humans from the crowd. 

Researchers from Bing have made a breakthrough in using LLMs to replace “the crowd” annotators (in whole or in part) in the 2023 paper, “Large language models can accurately predict searcher preferences.” 

The enormity of this recent work by Bing (in terms of the potential change for search research) was emphasized in a tweet by Sanderson. Sanderson described the talk as “incredible,” noting, “Synthetic labels have been a holy grail of retrieval research for decades.”

While sharing the paper and subsequent case study, Thomas also shared Bing is now using GPT-4 for its relevance judgments. So, not just research, but (to an unknown extent) in production search too.

Mark Sanderson on X

So what has Bing done?

The use of GPT-4 at Bing for relevance labeling

The traditional approach of relevance evaluation typically produces a varied mixture of gold and silver labels when “the crowd” provides judgments from explicit feedback after reading “the guidelines” (Bing’s equivalent of Google’s Quality Raters Guide). 

In addition, live tests in the wild utilizing implicit feedback typically generate gold labels (the reality of the real world “human in the loop”), but with a lack of scale and high relative costs. 

Bing’s approach utilized GPT-4 LLM machine-learned pseudo-relevance annotators created and trained via prompt engineering. The purpose of these instances is to emulate quality raters to detect relevance based on a carefully selected set of gold standard labels.

This was then rolled out to provide bulk “gold label” annotations more widely via machine learning, reportedly for a fraction of the relative cost of traditional approaches. 

The prompt included telling the system that it is a search quality rater whose purpose is to assess whether documents in a set of results are relevant to a query using a label reduced to a binary relevant / not relevant judgment for consistency and to minimize complexity in the research work.

To aggregate evaluations more broadly, Bing sometimes utilized up to five pseudo-relevance labelers via machine learning per prompt.

The approach and impacts for cost, scale and purported accuracy are illustrated below and compared with other traditional explicit feedback approaches, plus implicit online evaluation.

Interestingly, two co-authors are also co-authors in Bing’s research piece, “The Crowd is Made of People,” and undoubtedly are well aware of the challenges of using the crowd.

Source: “Large language models can accurately predict searcher preferences,” Thomas et al., 2023
Source: “Large language models can accurately predict searcher preferences,” Thomas et al., 2023

With these findings, Bing researchers claim:

“To measure agreement with real searchers needs high-quality “gold” labels, but with these we find that models produce better labels than third-party workers, for a fraction of the cost, and these labels let us train notably better rankers.” 

Scale and low-cost combined

These findings illustrate machine learning and large language models have the potential to reduce or eliminate bottlenecks in data labeling and, therefore, the evaluation process.

This is a sea-change pointing the way to an enormous step forward in how evaluation before algorithmic updates are undertaken since the potential for scale at a fraction of the cost of “the crowd” is considerable.

It’s not just Bing reporting on the success of machines over humans in relevance labeling tasks, and it’s not just ChatGPT either. Plenty of research into whether human assessors can be replaced in part or wholly by machines is certainly picking up pace in 2022 and 2023 in other research, too.

Others are reporting some success in utilizing machines over humans for relevance labeling, too

In a July 2023 paper, researchers at the University of Zurich found open source large language models (FLAN and HugginChat) outperform human crowd workers (including trained relevance annotators and consistently high-scoring crowd-sourced MTurk human relevance annotators). 

Although this work was carried out on tweet analysis rather than search results, their findings were that other open-source large language models were not only better than humans but were almost as good in their relevance labeling as ChatGPT (Alizadeh et al, 2023).

This opens the door to even more potential going forward for large-scale relevance annotations without the need for “the crowd” in its current format.

But what might come next, and what will become of ‘the crowd’ of human quality raters?

Responsible AI importance 

Caution is likely overwhelmingly front of mind for search engines. There are other highly important considerations.

Responsible AI, as yet unknown risk with these approaches, baked-in bias detection, and its removal, or at least an awareness and adjustment to bias, to name but a few. LLMs tend to “hallucinate,” and “overfitting” could present problems as well, so monitoring might well consider factors such as these with guardrails built as necessary. 

Explainable AI also calls for models to provide an explanation as to why a label or other type of output was deemed relevant, so this is another area where there will likely be further development. Researchers are also exploring ways to create bias awareness in LLM relevance judgments. 

Human relevance assessors are monitored continuously anyway, so continual monitoring is already a part of the evaluation process. However, one can presume Bing, and others, would tread much more cautiously with this machine-led approach over the “the crowd” approach. Careful monitoring will also be required to avoid drops in quality in exchange for scalability.

In outlining their approach (illustrated in the image above), Bing shared this process: 

  • Select via gold labels
  • Generate labels in bulk
  • Monitor with several methods

“Monitor with several methods” would certainly fit with a clear note of caution.

Next steps?

Bing, and others, will no doubt look to improve upon these new means of gathering annotations and relevance feedback at scale. The door is unlocked to a new agility.

A low-cost, hugely scalable relevance judgment process undoubtedly gives a strong competitive advantage when adjusting search results to meet changing information needs.

As the saying goes, the cat is out of the bag, and one could presume the research will continue to heat up to a frenzy in the information retrieval space (including other search engines) in the short to medium term.

A spectrum of human and machine assessors?

In their 2023 paper “HMC: A Spectrum of Human–Machine-Collaborative Relevance Judgement Frameworks,” Clarke et al. alluded to a feasible approach that might well mean subsequent stages of a move toward replacement of the crowd with machines taking a hybrid or spectrum form.

While a spectrum of human-machine collaboration might increase in favor of machine-learned methods as confidence grows and after careful monitoring, none of this means “the crowd” will leave entirely. The crowd may become much smaller, though, over time.

It seems unlikely that search engines (or IR research at large) would move completely away from using human relevance judges as a guardrail and a sobering sense-check or even to act as judges of the relevance labels generated by machines. Human quality raters also present a more robust means of combating “overfitting.”

Not all search areas are considered equal in terms of their potential impact on the life of searchers. Clarke et al., 2023, stress the importance of a more trusted human judgment in areas such as journalism, and this would fit well with our understanding as SEOs of Your Money or Your Life (YMYL).

The crowd might well just take on other roles depending upon the weighting in a spectrum, possibly moving into more of a supervisory role, or as an exam marker of machine-learned assessors, with exams provided for large language models requiring explanations as to how judgments were made.

Clarke et al. ask: “What weighting between human and LLMs and AI-assisted annotations is ideal?” 

What weighting of human to machine is implemented in any spectrum or hybrid approach might depend on how quickly the pace of research picks up. While not entirely comparable, if we look at the herd movement in the research space after the introduction of BERT and transformers, one can presume things will move very quickly indeed. 

Furthermore, there is also a massive move toward synthetic data already, so this “direction of travel” fits with that. 

According to Gartner:

  • “Solutions such as AI-specific data management, synthetic data and data labeling technologies, aim to solve many data challenges, including accessibility, volume, privacy, security, complexity and scope.” 
  • “By 2024, Gartner predicts 60% of data for AI will be synthetic to simulate reality, future scenarios and de-risk AI, up from 1% in 2021.” 

Will Google adopt these machine-led evaluation processes?

Given the sea-change to decades-old practices in the evaluation processes widely used by search engines, it would seem unlikely Google would not at least be looking into this very closely or even be striving towards this already. 

If the evaluation process has a bottleneck removed via the use of large language models, leading to massively reduced data sparsity for relevance labeling and algorithmic update feedback at lower costs for the same, and the potential for higher quality levels of evaluation too, there is a certain sense in “going there.”

Bing has a significant commercial advantage with this breakthrough, and Google has to stay in and lead, the AI game.

Removals of bottlenecks have the potential to massively increase scale, particularly in non-English languages and into additional markets where labeling might have been more difficult to obtain (for example, the subject matter expert areas or the nuanced queries around more technical topics). 

While we know that Google’s Search Generative Experience Beta, despite expanding to 120 countries, is still considered an experiment to learn how people might interact with or find useful, generative AI search experiences, they have already stepped over the “AI line.”

Greg Gifford on X - SGE is an experiment

However, Google is still incredibly cautious about using AI in production search.

Who can blame them for all the antitrust and legal cases, plus the prospect of reputational damage and increasing legislation related to user privacy and data protection regulations?

James Manyika, Google’s senior vice president of technology and society, speaking at Fortune’s Brainstorm AI conference in December 2022, explained:

“These technologies come with an extraordinary range of risks and challenges.” 

However, Google is not shy about undertaking research into the use of large language models. Heck, BERT came from Google in the first place. 

Certainly, Google is exploring the potential use of synthetic query generation for relevance prediction, too. Illustrated in this recent 2023 paper by Google researchers and presented at the SIGIR information retrieval conference.

Google paper 2023 on relevance prediction

Since synthetic data in AI/ML reduces other risks that might relate to privacy, security, and the use of user data, simply generating data out of thin air for relevance prediction evaluations may actually be less risky than some of the current practices.

Add to the other factors that could build a case for Google jumping on board with these new machine-driven evaluation processes (to any extent, even if the spectrum is mostly human to begin with):

  • The research in this space is heating up. 
  • Bing is running with some commercial implementation of machine over people labeling. 
  • SGE needs loads of labels.
  • There are scale challenges with the status quo.
  • The increasing spotlight on the use of low-paid workers in the data-labeling industry overall. 
  • Respected information retrieval researchers are asking is now the time to revisit the use of machines over humans in labeling?

Openly discussing evaluation as part of the update process

Google also seems to be talking much more openly of late about “evaluation” too, and how experiments and updates are undertaken following “rigorous testing.” There does seem to be a shift toward opening up the conversation with the wider community.

Here’s Danny Sullivan just last week giving an update on updates and “rigorous testing.”

Martin Splitt on X - Search Central Live

And again, explaining why Google does updates.

Greg Bernhardt on X

Search off The Record recently discussed “Steve,” an imaginary search engine, and how updates to Steve might be implemented based on the judgments of human evaluators, with potential for bias, amongst other points discussed. There was a good amount of discussion around how changes to Steve’s features were tested and so forth. 

This all seems to indicate a shift around evaluation unless I am simply imagining this.

In any event, there are already elements of machine learning in the relevance evaluation process, albeit implicit feedback. Indeed, Google recently updated its documentation on “how search works” around detecting relevant content via aggregated and anonymized user interactions.

“We transform that data into signals that help our machine-learned systems better estimate relevance.”

So perhaps following Bing’s lead is not that far a leap to take after all?

What if Google takes this approach?

What might we expect to see if Google embraces a more scalable approach to the evaluation process (huge access to more labels, potentially with higher quality, at lower cost)?

Scale, more scale, agility, and updates

Scale in the evaluation process and speedy iteration of relevance feedback and evaluations pave the way for a much greater frequency of updates, and into many languages and markets.

An evolving, iterative, alignment with true relevance, and algorithmic updates to meet this, could be ahead of us, with less broad sweeping impacts. A more agile approach overall. 

Bing takes a much more agile approach in their evaluation process already, and the breakthrough with LLM as relevance labeler makes them even more so. 

Fabrice Canel of Bing, in a recent interview, reminded us of the search engine’s constantly evolving evaluation approach where the push out of changes is not as broad sweeping and disruptive as Google’s broad core update or “big” updates. Apparently, at Bing, engineers can ideate, gain feedback quickly, and sometimes roll out changes in as little as a day or so.

All search engines will have compliance and strict review processes, which cannot be conducive to agility and will no doubt build up to a form of process debt over time as organizations age and grow. However, if the relevance evaluation process can be shortened dramatically while largely maintaining quality, this takes away at least one big blocker to algorithmic change management.

We have already seen a big increase in the number of updates this year, with three broad core updates (relevance re-evaluations at scale) between August and November and many other changes concerning spam, helpful content, and reviews in between.

Coincidentally (or probably not), we’re told “to buckle up” because major changes are coming to search. Changes designed to improve relevance and user satisfaction. All the things the crowd traditionally provides relevant feedback on.

Kenichi Suzuki on X

So, buckle up. It’s going to be an interesting ride.

rustybrick on X - Google buckle up

If Google takes this route (using machine labeling in favor of the less agile “crowd” approach), expect a lot more updates overall, and likely, many of these updates will be unannounced, too. 

We could potentially see an increased broad core update cadence with reduced impacts as agile rolling feedback helps to continually tune “relevance” and “quality” in a faster cycle of Learning to Rank, adjustment, evaluation and rollout.

Gianluca Fiorelli on X - endless updates

The post Quality rater and algorithmic evaluation systems: Are major changes coming? appeared first on Search Engine Land.

Original source: https://searchengineland.com/quality-rater-algorithmic-evaluation-systems-changes-434895

4 B2B trends that will catapult you ahead of the competition next year by Cynthia Ramsaran

Aircraft catapult

More frequently than not, B2B buyers are vetting sellers based not only on product specifications, pricing and other traditional factors but on the digital experiences they deliver. Failing to adapt to these rising customer expectations can be costly. 

Deloitte Digital conducted a study of more than 500 B2B executives at U.S. companies and discovered that 77% of B2B executives agree that digital transformation is critical to their company’s success.

Join experts from Deloitte Digital, who unveil the research findings and highlight the four trends that lead to stronger customer relationships all around: higher satisfaction, stronger spending, better retention and deeper trust.

Learn more by registering and attending “4 B2B Selling Trends to Catapult You Ahead of the Competition,” presented by Deloitte.


Click here to view more Search Engine Land webinars.

The post 4 B2B trends that will catapult you ahead of the competition next year appeared first on Search Engine Land.

Original source: https://searchengineland.com/4-b2b-trends-that-will-catapult-you-ahead-of-the-competition-next-year-434959

Native App Development vs. Hybrid App Development

Home Business Magazine Online

Nowadays, mobile apps are a crucial component of customer engagement as well as encouraging the company’s staff and partners. What to choose: native, cross-platform, or hybrid app development has become the subject of the heated debate among mobile app developers. Still, there’s no clear-cut answer as each situation is case-based.

When it comes to developing mobile applications, the choice between native and hybrid approaches is pivotal. Each methodology presents distinct advantages and trade-offs. This comparison aims to delineate the disparities between native and hybrid app development, elucidating the key factors that businesses and developers should consider in making informed decisions for their mobile projects.

Why Choose Native App Development?

Choosing native app development has several advantages, making it a preferred option for many projects. Here are some reasons why one might opt for native app development:

1. Performance:

Native apps are developed specifically for a particular platform (iOS or Android), allowing to leverage the full potential of the device’s hardware and software. This often results in superior performance compared to hybrid or cross-platform alternatives.

2. User Experience:

Native apps offer a seamless and consistent user experience, as they adhere to the design guidelines and UI patterns of the respective platforms (Material Design for Android, Human Interface Guidelines for iOS). This leads to intuitive navigation and familiarity for users.

3. Access to Device Features:

Native development provides direct access to the device’s features and capabilities, such as the camera, GPS, accelerometer, and more. This enables developers to create functionality-rich apps that can take full advantage of the hardware.

4. Optimized for Platform Updates:

Native apps are typically quicker to adopt new features and updates introduced by operating systems. This ensures that the app remains compatible, performs well, and takes advantage of the latest functionalities.

5. Better Security:

Native apps benefit from the security features of the underlying operating system. Additionally, they are subject to stricter app review processes on app stores, reducing the likelihood of malicious software.

6. Offline Functionality:

Native apps can often offer more robust offline capabilities, allowing users to access certain features or content even when they are not connected to the internet.

7. Community and Support:

Both iOS and Android platforms have large and active developer communities. Choosing native development means tapping into these communities for support, libraries, and resources.

8. App Store Optimization (ASO):

Native apps tend to perform better in app store rankings and visibility. Platforms like the Apple App Store and Google Play Store often prioritize native apps, potentially leading to better discoverability.

9. Customization:

Native development allows for high customization, making it easier to tailor the app’s user interface and user experience to match the specific requirements and preferences of the platform and target audience.

10. Reliability and Stability:

Native apps are generally more reliable and stable because they are optimized for the specific platform. This can result in fewer crashes and better overall performance.

While native app development offers these advantages, it’s important to note that it may require separate development efforts for iOS and Android versions, potentially increasing costs and development time. The choice between native depends on factors such as project requirements, budget, and the target audience.

Why Choose Hybrid App Development?

Choosing hybrid app development comes with its own set of advantages, making it a suitable option for certain projects. Here are reasons why one might opt for hybrid app development:

1. Cross-Platform Compatibility:

One of the primary advantages of hybrid development is the ability to write code once and deploy it across multiple platforms (iOS, Android, etc.). This can significantly reduce development time and costs compared to building separate native apps.

2. Cost-Effectiveness:

Hybrid apps can be more cost-effective than native development, as a single codebase can be used for multiple platforms. This is particularly beneficial for businesses with budget constraints or those looking to optimize development resources.

3. Faster Development Time:

Since a single codebase can be reused for different platforms, development time is often shorter for hybrid apps compared to building separate native versions. This can lead to a quicker time-to-market.

4. Easier Maintenance:

Updates and maintenance are streamlined in hybrid development since changes made to the codebase apply to both platforms simultaneously. This makes it easier to roll out bug fixes and new features.

5. Web Technology Stack:

Hybrid apps are typically built using web technologies such as HTML, CSS, and JavaScript. Developers with expertise in web development can leverage their skills to create mobile apps, potentially reducing the need for additional specialized knowledge.

6. Access to Device Features:

Hybrid apps can access certain device features through plugins or frameworks like Apache Cordova and React Native. While not as direct as native access, it allows for a balance between cross-platform development and device capabilities.

7. Uniform User Experience:

Hybrid apps aim to provide a consistent user experience across platforms, maintaining a similar look and feel. This can be advantageous for brand consistency and ensuring a familiar interface for users.

8. Offline Functionality:

Hybrid apps can incorporate offline capabilities, allowing users to access certain features or content without an internet connection. This is achieved through technologies like local storage and caching.

9. Simplified Deployment:

Hybrid apps can be deployed through a single codebase, simplifying the deployment process. Updates and changes can be pushed to both platforms simultaneously, reducing the complexity of the release process.

10. Broader Audience Reach:

Hybrid apps enable businesses to reach a broader audience by offering a consistent experience on both major mobile platforms. This can be particularly beneficial for startups or companies targeting diverse user bases.

11. Adaptability to Business Changes:

Hybrid apps offer flexibility and adaptability, making it easier for businesses to pivot or modify their apps based on changing requirements or feedback without undergoing a complete redevelopment.

While hybrid app development offers these advantages, it’s crucial to consider the specific requirements of the project, including performance needs, access to native features, and user experience expectations. The choice between hybrid and other development approaches depends on factors such as project goals, budget constraints, and the desired balance between cross-platform compatibility and native capabilities.

Are Hybrid Apps the Future?

As mobile apps gain popularity and the demand for mobile app development services increases, companies are looking for different methods to create top-notch apps for clients. The hybrid approach is very appealing to many customers.

Indeed, it all comes from the project requirements. For example, if the company is releasing a new product or service and yet it’s not clear whether people need that, a hybrid approach will come in handy to test this hypothesis. Another example, in case it deals with displaying content and there’s no anticipated interaction with the interface, switching between screens or tabs, and so forth, hybrid app development is a good match.

Conclusion

In conclusion, the choice between native and hybrid app development hinges on a variety of factors, each approach presenting distinct advantages and trade-offs. Native development offers unparalleled performance, seamless access to device features, and a superior user experience tailored to specific platforms. On the other hand, hybrid development offers cost-effectiveness, quicker time-to-market, and the efficiency of a single codebase across multiple platforms.

Ultimately, the decision should be guided by the unique requirements of the project, the target audience, and the overarching business goals. Native development excels in scenarios where optimal performance and full utilization of platform-specific capabilities are critical. Conversely, hybrid development is a compelling choice when cross-platform compatibility, cost efficiency, and faster development cycles take precedence.

Whether opting for native or hybrid development, the key lies in understanding the project’s nuances, aligning the chosen approach with business objectives, and delivering a mobile app that meets the needs and expectations of users. As technology continues to evolve, staying informed about the latest trends and advancements in both native and hybrid development will empower decision-makers to make strategic choices that best suit their specific contexts.

The post Native App Development vs. Hybrid App Development appeared first on Home Business Magazine.

Original source: https://homebusinessmag.com/businesses/app-development/native-app-development-vs-hybrid-app-development/

Fresh Approaches to Real Estate Investing

Home Business Magazine Online

Traditional real estate investment is well-established and offers several advantages. But whether you’re just starting out or have been in the field for years, you might be searching for new, unconventional ways to earn passive income. With the rise of services like NewHomesMate, which simplify the process of finding new homes for sale, investors now have more options at their fingertips.

This article will cover three strategies that can enhance your real estate portfolio through creative investing. So keep reading if you want to learn more about this topic.

House Hacking

House hacking is all about buying a property with several units, living in one, and renting out the rest. The idea is that the rent from your tenants covers your mortgage, meaning you live virtually rent-free. This strategy is gaining popularity, especially for those looking to manage their mortgage payments, property taxes, and maintenance costs. Plus, it’s a great opportunity to build equity in your home at the same time.

Self-Storage Unit Investment

Investing in self-storage units is a unique choice in the real estate world, and it’s getting popular really quickly. Mainly, that’s because it is much simpler than the typical landlord responsibilities like ongoing maintenance or tenant management.

The process involves either purchasing an existing storage facility or building a new one. Once it’s up and running the right property management software can do most of the heavy lifting for you. Imagine managing your investment with just a few taps on your phone!

Here is how it works:

  1. A customer arrives and selects a storage unit;
  2. They sign the lease agreement;
  3. They fill the unit with their belongings.

And that’s the entire process. You’re not constantly chasing rent or handling late-night maintenance issues. This approach to real estate investment is all about minimal hassle and stress.

House Flipping

Imagine taking a house that’s seen better days and turning it into someone’s dream home. That’s the essence of house flipping. Often known as “fix and flip,” this strategy is about buying a property at a lower value, revamping it, and then selling it for a higher price. It’s not just a hobby for many investors; house flipping can be a highly profitable full-time job when done right.

And the most appealing part is that, unlike regular rental management, house flipping is more about the transformation than the long-term care. You’re not entangled in the complexities of tenant management or ongoing maintenance.

Thinking Differently in Real Estate Investment

These new ways of investing in real estate show us that there are many different paths to success. From making over houses in house flipping to the easy management of self-storage units or living almost rent-free through house hacking, each method has its own unique benefits. They’re suited for different kinds of people and can be a great way to make money and grow in the real estate world. As the market changes, so do the ways we can get involved, making it an exciting time to try new things in real estate.

The post Fresh Approaches to Real Estate Investing appeared first on Home Business Magazine.

Original source: https://homebusinessmag.com/businesses/real-estate/fresh-approaches-real-estate-investing/