Supply Chain

Greenfield Network Design
Where Would You Put the DCs If You Started Over?

Greenfield analysis is the clean-slate first pass of network design. It places facilities where demand clusters, fast and in a way anyone can follow, and then hands off to the harder facility-selection work.

What problem does it solve?

Greenfield analysis answers a narrow but very common question: “If we were starting from scratch, with no existing facilities and no historical commitments, where on the map would we put N distribution centers to be closest to demand?”

It’s a clean-slate analysis. The output is a set of locations where facilities should go, given where customers are and how much they buy. It doesn’t consider inbound freight, fixed operating costs, capacity tiers or supplier locations. It’s deliberately simpler than full network optimization, and that simplicity is the feature: greenfield gives you a fast, defensible starting answer that you then refine.

When to use it, and when not to

Reach for greenfield when:

Don’t use greenfield when:

A well-run network design project usually starts with greenfield (minutes of computation and an easy story for executives), then moves to facility selection, the harder MIP that includes the existing footprint and costs, once the candidates are on the table.

The math: center of gravity

The classic greenfield algorithm is the weighted Euclidean center of gravity. For a single facility serving a set of customers, you want the point (x, y) that minimizes:

Weighted distance to demand

minimize  Σ_j  d_j · sqrt( (x - x_j)² + (y - y_j)² )

where:

This is sometimes called the Weber problem (Alfred Weber, 1909). It answers “where should the warehouse go to minimize average weighted travel distance?” Center of gravity analysis walks through the method in plain terms, with a worked example.

There’s no closed-form solution, but a simple iterative algorithm converges quickly. Start at the demand-weighted centroid:

Starting point

x₀ = Σ d_j · x_j  /  Σ d_j
y₀ = Σ d_j · y_j  /  Σ d_j

Then iterate:

Weiszfeld update

x_{n+1} = Σ (d_j · x_j / dist_nj)  /  Σ (d_j / dist_nj)
y_{n+1} = Σ (d_j · y_j / dist_nj)  /  Σ (d_j / dist_nj)

Here dist_nj is the distance from the current point to customer j. It converges in about 5 to 20 iterations for well-behaved demand maps. This is Weiszfeld’s algorithm, a fixed-point iteration on the gradient of the objective.

More than one facility: location-allocation

When you want N facilities rather than one, the problem splits in two: assign each customer to one facility, and place each facility at the weighted centroid of its assigned customers. The two decisions are coupled. Change the assignments and the best locations shift; move the locations and the best assignments shift.

The standard algorithm alternates:

  1. Assign. Given the current facility locations, assign each customer to its nearest facility, by straight-line distance, weighted distance or road distance if you have it.
  2. Locate. Given the assignments, move each facility to the weighted centroid of its customers, using the single-facility math above.
  3. Repeat until the assignments stop changing.

This is the same structure as Lloyd’s algorithm for k-means clustering, the one used in machine learning to cluster data points. The only difference is demand-weighted centroids, so a high-volume customer pulls the facility toward it.

The algorithm is greedy and can get stuck in a local optimum. Practical implementations run several random starts and keep the best result.

Service constraints: share served within a distance

In practice, the question isn’t just “minimize average weighted distance.” It’s “how many facilities do I need so that 95% of demand is within 500 miles of some facility?”

That adds a coverage constraint on top of location-allocation:

The two give different answers when demand is uneven. Share of demand favors big customers, so a 500,000-unit customer outweighs a 5,000-unit one. Share of customers counts every site equally. Real network design conversations usually need both.

With service constraints, the question becomes “what’s the smallest N that meets the constraint?” The solver tries N = 1 and checks service, tries N = 2, and stops at the smallest N that passes.

Why results snap to the nearest city

The algorithm minimizes weighted distance to a precise latitude and longitude, but that mathematical optimum can land in a cornfield, a lake or a restricted industrial zone. The Sandbox’s greenfield engine always snaps each result to the nearest city instead of reporting the raw centroid. There are three reasons:

The tradeoff is that the nearest-city result sits a few miles from the unconstrained optimum. For strategic planning, which is what greenfield is for, that’s noise. If you need the exact centroid, for example to compare against a candidate site you’re already evaluating, the underlying algorithm has it. The public demo just doesn’t show it.

Limitations

Greenfield is fast and easy to explain, but it assumes a lot:

Think of greenfield as a starting argument, not a final answer: “These N centroids are where customers cluster. Now let’s run facility selection to see which existing or candidate facilities best match them, while honoring everything greenfield ignores.”

Greenfield, brownfield and footprint rationalization

Network design questions come in three shapes, and they use different methods. The Sandbox has a demo of each.

Greenfield: start from demand

Greenfield ignores the current footprint and places facilities where demand clusters, using the methods above. It’s the right question for a new entrant, or for a “what if we started over?” baseline. The Greenfield · US demo sites two to eight DCs from scratch for 189 US demand points.

Brownfield: choose which existing sites to keep

Brownfield starts from the sites you have. No new locations are sited. Some DCs are pinned open because of a lease or a key customer, the rest are optional, and the engine chooses which optional sites to close to reach a target count. The Brownfield · EU demo has 142 demand points in France, Germany and Luxembourg and ten DCs. Its graph shows what your pins cost against a run where every site is optional.

Footprint rationalization: consolidate a network that grew

Footprint rationalization is brownfield on a bigger, older network. It starts from today’s footprint and asks how far you can consolidate, and which DCs go first. The Footprint rationalization demo opens on a North American network of 658 customer points and sixteen DCs. It also shows the saving from re-assigning customers to their best DC before anything closes.

All three demos score demand times distance, not landed cost, so treat them as the first pass. Once fixed costs, capacity and freight rates matter, the question becomes facility-selection optimization. The count itself is covered in how many distribution centers do you need?

Try it in the Sandbox

The Greenfield Design, US demo runs on a 189-customer US sample. Step the DC count from 2 to 8 and watch the elbow in the cost-vs-DC-count curve, which is where a network design conversation usually starts. The Brownfield Design, EU demo shows the next step: pin some sites open and let the engine choose which optional sites to keep.

For a greenfield analysis on your own demand data, including service-area constraints, multi-product weighting and the facility-selection work that follows, talk to us.

Sources: archived enterprise greenfield-analysis project documentation and network optimization course material (2013–2015); standard site-selection methodology; the Weber location problem (Alfred Weber, 1909); Weiszfeld’s algorithm; Lloyd’s algorithm for k-means clustering.

Site 2 to 8 DCs on a US demand map

The Greenfield demo runs in your browser, with nothing to install.

Open Greenfield Design →

Or read how network optimization works.